Description
NVIDIA is seeking an NCX Senior Engineer to join the DSX team, focusing on NVIDIA Cloud Partner (NCP) infrastructure operations.
The role involves guiding partners beyond initial cluster deployment and validation into advanced Day 2 operations, covering infrastructure health, observability, lifecycle management, and performance validation.
Responsibilities:
- Lead NCP Day 2 operational readiness efforts, collaborating with partners to set up systems, procedures, automation, and operational methods.
- Develop and implement continuous infrastructure validation methods for large-scale AI clusters.
- Establish observability and operational telemetry across compute, GPU, networking, storage, Kubernetes, and AI workloads.
- Build automated detection and remediation workflows to minimize disruption to customer workloads.
- Implement scalable strategies for managing sizable GPU fleets, including driver and firmware lifecycle administration.
- Operationalize NVIDIA reference architectures, translating requirements into production operating practices.
- Define operational health and readiness, developing health signals, SLOs, and validation mechanisms.
- Build reusable operational frameworks, including tooling, automation, and reference implementations.
Requirements:
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field.
- 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, or similar roles.
- Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
- Deep understanding of Kubernetes, containers, cluster scheduling, and production observability.
- Experience with automation for infrastructure lifecycle management and failure detection.
- Strong networking fundamentals and experience troubleshooting complex distributed systems.
- Programming and automation experience using Python, Go, shell scripting, or similar languages.
Preferred qualifications:
- Experience managing extensive GPU or accelerated computing infrastructure for AI workloads.
- Familiarity with NVIDIA technologies, including DGX/HGX systems, CUDA, and NVLink/NVSwitch.
- Proven experience collaborating with NVIDIA Cloud Partners and hyperscale cloud providers.
- Extensive knowledge of infrastructure observability tools, including Prometheus, Grafana, and OpenTelemetry.
Benefits:
- Equity eligibility
- Comprehensive benefits package
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/NCX-Senior-Engineer_JR2024060