# Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--Distributed-Systems-Engineer---DGX-Cloud_JR2021856-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_80232e49-a0a

## Description

NVIDIA is seeking experienced software engineers to help scale its AI Infrastructure. As a Senior Software Engineer, Distributed Systems Engineer on the DGX Cloud team, you will be responsible for production systems that enable large scalable GPU clusters for various AI workloads.

You will design and develop a massively distributed scalable platform to identify, diagnose, and remediate non-performant GPU assets. You will work with teams across NVIDIA to ensure production AI clusters run reliably and consistently with maximum performance, evaluating system failures and improving services based on a well-defined incident management process.

To succeed in this role, you should have:

- Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work.

- Strong communication skills, with the ability to work successfully with multi-functional teams, principles, and architects.

- 5+ years of experience in a similar role, with experience on large-scale production systems.

- A BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience.

- Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.

Preferred qualifications include:

- Technical competency in managing and automating large-scale distributed systems independent of cloud providers.

- Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Base Command Manager).

- Prior experience in asynchronous workflows and/or event-driven architecture.

- Proven operational excellence in maintaining reliable and performant infrastructure.

NVIDIA offers equity and benefits to its employees.

## Skills

### Required
- software engineering
- cluster operations
- operator development
- node health monitoring
- GPU resource scheduling
- Go
- Python
- data structures
- algorithms

### Nice to have
- Kubernetes
- Slurm
- Base Command Manager
- asynchronous workflows
- event-driven architecture

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--Distributed-Systems-Engineer---DGX-Cloud_JR2021856-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
