# Senior Software Engineer, DGX Cloud Orchestration

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--DGX-Cloud-Orchestration_JR2020967-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9c8fa9b3-ab5

## Description

We are looking for a Senior Software Engineer to join our DGX Cloud team and build foundational systems that drive NVIDIA's high-performance GPU infrastructure.

You will play a critical role in designing scalable automation solutions, integrating diverse systems, and enabling seamless workflows across global cloud operations.

**Responsibilities:**

- Design and develop APIs to orchestrate and integrate operational workflows.

- Build state management and workflow automation systems that streamline infrastructure lifecycle processes.

- Collaborate across teams to codify business processes into scalable, self-measuring systems.

- Develop extensible, schema-driven platforms for reducing manual toil and ensuring operational consistency.

- Drive integrations with container orchestration tools like Kubernetes and observability systems such as Prometheus, OpenTelemetry, Grafana.

- Optimize the reliability and efficiency of cloud operations through automated workflows and telemetry systems.

- Lead and ship impactful technical projects, ensuring quality and scalability at every stage.

**Requirements:**

- 8+ years of industry experience with a Bachelor's or Master's degree (or equivalent experience), or 2+ years with a PhD.

- Expertise in designing, building, and operating services in a high reliability environment.

- Proficiency in programming languages such as Go, Java, or Python.

- Strong understanding of cloud infrastructure (AWS, GCP, Azure) and container technologies like Docker and Kubernetes.

- Experience with high-scale distributed systems, including architectural patterns for APIs and data pipelines.

- Outstanding communication and collaboration skills, with a focus on solving complex operational challenges.

- A passion for automating manual processes and driving system efficiency.

**Benefits:**

- Equity

- Benefits package

## Skills

### Required
- cloud infrastructure
- container technologies
- distributed systems
- API design
- automation
- Kubernetes
- Prometheus
- OpenTelemetry
- Grafana
- Go
- Java
- Python

### Nice to have
- workflow orchestration systems
- operational aspects of NVIDIA AI/ML software stack
- debugging and problem-solving skills in distributed environments

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer--DGX-Cloud-Orchestration_JR2020967-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
