# Senior Software Engineer - AI Research Clusters

**Company**: NVIDIA
**Location**: Santa Clara
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer---AI-Research-Clusters_JR2013615?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_37ec207c-dc8

## Description

We are seeking a Senior Software Engineer to join our AI Research Clusters team. In this role, you will propose and implement engineering solutions to ensure delivery of functional, reliable, secure, and performance-optimal GPU clusters to internal researchers. Your work will empower scientists and engineers to train, fine-tune, and deploy the most advanced ML models on some of the world's most powerful GPU systems.

**Responsibilities:**

- Work with coworkers across the AI Platform organization to understand the pain points of validating, monitoring and operating GPU clusters at scale.

- Design, develop and maintain engineering solutions to solve those pain points systematically.

- Research in traditional AIOps and the emerging Agentic AI, and leverage it to further reduce the operation toil.

- Participate in on-call support for systems, platforms built and owned by the team.

**Requirements:**

- BS/MS in Computer Science, Engineering, or equivalent experience.

- 5+ years in software/platform engineering, including 3+ years in ML infrastructure or distributed systems.

- Experience in software development lifecycle on Linux-based platforms.

- Strong coding skills in languages such as Python, C++ or Rust.

- Experience with Docker, Kubernetes, GitLab CI, automated deployments.

- Experience with AIOps or Agentic AI and apply it successfully in production environment.

**Nice to Have:**

- Proficiency with full-stack development: Relational Data Modeling, DB optimization, REST API Semantics, Javascript, CSS, providing API as a service.

- Passion for building developer-centric platforms with great UX and strong operational reliability.

- Experience running Slurm or custom scheduling frameworks in production ML environments.

- Familiarity with GPU computing, Linux systems internals, and performance tuning at scale.

## Skills

### Required
- Python
- C++
- Rust
- Docker
- Kubernetes
- GitLab CI
- AIOps
- Agentic AI

### Nice to have
- Relational Data Modeling
- DB optimization
- REST API Semantics
- Javascript
- CSS
- Slurm
- GPU computing
- Linux systems internals
- performance tuning

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer---AI-Research-Clusters_JR2013615?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
