Description
NVIDIA is hiring a Senior Engineer to join their DSX team, working closely with strategic NVIDIA Cloud Partners to build and improve operational capabilities for large-scale NVIDIA accelerated infrastructure. The role involves guiding partners beyond initial cluster deployment and validation into advanced Day 2 operations.
Responsibilities:
- Lead NCP Day 2 operational readiness efforts
- Build continuous infrastructure validation
- Establish observability and operational telemetry
- Develop automated detection and remediation
- Refine fleet lifecycle administration
- Operationalize NVIDIA reference architectures
- Define operational health and readiness
- Build reusable operational frameworks
Requirements:
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field
- 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles
- Strong experience operating Linux-based distributed systems and cloud infrastructure in production
- Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments
- Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements
- Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management
- Strong networking fundamentals and experience troubleshooting complex distributed systems
- Programming and automation experience using Python, Go, shell scripting, or similar languages
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Germany-Remote/Senior-Engineer--NCX_JR2024831