Nvidia
Technical Architect
Overview
We are looking for a strong technical architect to own end-to-end system software architecture for Space-1 and successor orbital platforms. You will architect the full stack — application to libraries, from data center stack to BMC and BIOS firmware, manageability, and telemetry through the host OS, GPU and CPU drivers, and CUDA — to deliver a production-ready inference platform that operates reliably in the radiation, thermal-cycling, and remote-operations environment of LEO.
About Nvidia
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and un
Requirements & Eligibility
- 15+ years of relevant experience in server/platform system software — spanning compute libraries, BMC firmware, BIOS, host OS, drivers, and manageability
- BS, MS, or PhD in EE/CS or related field of education (or equivalent experience)
- Working experience in building AI infrastructure and systems in space
- Strong knowledge of server architecture, data center manageability, and full-stack integration of firmware with OS and accelerator software
- Solid understanding of hardware management interfaces (USB, SMBus/I2C, PCIe) and proficiency with modern management protocols including Redfish, MCTP, and PLDM
- Strong and demonstrable skill in C/C++ and Python
Key Responsibilities
- Own system architecture for inference stack and other applications running on this class of products and make it resilient to any fault happening in space
- Co-architect with the orbital hardware system architecture team to define interfaces, partitioning, and trade-offs across silicon, board, firmware, OS, and AI workload layers for 5-year LEO missions
- Own end-to-end system software architecture for Space-1 and successor Orbital Data Center modules
- Define the manageability architecture for an unreachable, autonomous data center
- Architect rad-tolerant system software behaviors — ECC handling, memory scrubbing, latch-up mitigation, deterministic recovery, and graceful degradation through 5 years and up to ~8,000 thermal cycles in dawn–dusk sun-synchronous orbit
- Drive Redfish, MCTP, PLDM, and constellation-level management protocols across BMC, BIOS, and host software
- Define minimum BMC feature set, pin budget, boot architecture, and dual-module redundancy strategy
- Partner with cloud and constellation customers to translate mission requirements into actionable platform software architecture
Disclaimer: Trace Hiring is an independent job board. We are not directly affiliated with Nvidia. Please verify all details on the official company application portal.