Description
We are hiring experienced Senior Production Engineers to help scale up our AI Infrastructure. As a Senior Production Engineer, you will be part of the DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads.
Your key responsibilities will include: Implementing monitoring and health management capabilities that enable industry-leading reliability, availability, and scalability of GPU assets. Working with teams across NVIDIA to ensure production AI clusters run reliably and consistently with maximum performance. Evaluating system failures and improving services based on a well-defined incident management process.
To succeed in this role, you will need: Direct experience in a Production Engineering/DevOps/SRE role within a highly technical organization with demonstrable impact from your work. Strong communication skills, able to work successfully with multi-functional teams, principles, and architects. 8+ years in similar role and experience on large-scale production systems. Experience with Production Engineering/DevOps/SRE principles, tools, and techniques. A BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience. Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.
If you have technical competency in managing and automating large-scale distributed systems independent of cloud providers, advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Bright Cluster Manager), and proven operational excellence in maintaining reliable and performant AI infrastructure, you will stand out from the crowd.
You will also be eligible for equity and benefits.