Description
NVIDIA is seeking a principal-level software engineer to build the next generation of our Kubernetes platform. Our teams build foundational capabilities for self-service GPU infrastructure, managed Kubernetes control planes, management of cluster operations, and automation for large-scale AI and improved computational environments.
In this role, you will work at the intersection of platform architecture and production-grade Kubernetes lifecycle systems. You will lead technical strategy and execution for systems that make clusters easier to provision, upgrade, operate, and scale across cloud and on-premises environments.
Responsibilities:
- Lead the architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations.
- Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
- Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack.
- Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
- Lead the diagnosis and resolution of complex platform issues spanning infrastructure, runtime, networking, hardware, and operations, improving the scalability, resilience, and operability of systems supporting large-scale AI deployments.
- Influence engineering standards, architectural decisions, and long-term platform strategy.
- Mentor senior engineers and raise the bar for design quality, execution, and engineering rigor across the organization.
Requirements:
- BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 15+ years of relevant software engineering experience, including experience building and operating large-scale production systems.
- Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
- A strong background in distributed systems design, reliability, scalability, and failure recovery.
- Proven experience building platform software, infrastructure control planes, or foundations for managed services.
- Strong programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++.
- Experience designing clear APIs and abstractions for platform consumers and engineering teams.
- Demonstrated ability to provide technical leadership across team boundaries and drive ambiguous, cross-functional initiatives to completion.
- Excellent communication and collaboration skills, backed by a sustained record of significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives.
Benefits:
- Competitive salaries
- Generous benefits package
- Equity
Applications for this job will be accepted at least until August 15, 2026.