Description
As a Principal Rack Scale Systems Infrastructure Engineer at NVIDIA, you will build and guide the development of software systems that support our upcoming rack-scale infrastructure products and services. These systems sit where software meets hardware, and you will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.
Your responsibilities will include defining the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. You will use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate, including controllers, operators, reconciliation loops, and open source components. These components should operate safely at rack and fleet scale. You will also bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces.
You will partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. You will establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.
You will mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development. You will make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.
We are looking for someone with a strong background in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering. You should have solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs. You should also have practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.
You should have experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. You should also have experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.
You should have a strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. You should also have expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols.
You should be able to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation. You should also be able to craft software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.
You should have experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation. You should also have established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations.
You should have strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.
You will also be eligible for equity and benefits.