# Senior Engineer, NCX

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Germany-Remote/Senior-Engineer--NCX_JR2024831?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_bf5c7b37-071

## Description

NVIDIA is hiring a Senior Engineer to join their DSX team, working closely with strategic NVIDIA Cloud Partners to build and improve operational capabilities for large-scale NVIDIA accelerated infrastructure. The role involves guiding partners beyond initial cluster deployment and validation into advanced Day 2 operations.

Responsibilities:

- Lead NCP Day 2 operational readiness efforts

- Build continuous infrastructure validation

- Establish observability and operational telemetry

- Develop automated detection and remediation

- Refine fleet lifecycle administration

- Operationalize NVIDIA reference architectures

- Define operational health and readiness

- Build reusable operational frameworks

Requirements:

- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field

- 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles

- Strong experience operating Linux-based distributed systems and cloud infrastructure in production

- Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments

- Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements

- Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management

- Strong networking fundamentals and experience troubleshooting complex distributed systems

- Programming and automation experience using Python, Go, shell scripting, or similar languages

## Skills

### Required
- Linux
- Kubernetes
- containers
- cloud infrastructure
- distributed systems
- Python
- Go
- shell scripting
- automation
- infrastructure lifecycle management
- failure detection
- remediation
- upgrades
- configuration management
- networking
- troubleshooting

### Nice to have
- GPU
- accelerated computing
- AI training
- inference workloads
- NVIDIA technologies
- DGX/HGX systems
- CUDA
- NVLink/NVSwitch
- NVIDIA networking
- InfiniBand
- RoCE
- GPU Operator
- Network Operator
- Prometheus
- Grafana
- OpenTelemetry
- Alertmanager

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Germany-Remote/Senior-Engineer--NCX_JR2024831?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
