New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Software Engineer - AI Research Clusters

NVIDIA
Apply →
onsite senior full-time Santa Clara

First indexed 18 May 2026

Description

We are seeking a Senior Software Engineer to join our AI Research Clusters team. In this role, you will propose and implement engineering solutions to ensure delivery of functional, reliable, secure, and performance-optimal GPU clusters to internal researchers. Your work will empower scientists and engineers to train, fine-tune, and deploy the most advanced ML models on some of the world's most powerful GPU systems.

Responsibilities:

  • Work with coworkers across the AI Platform organization to understand the pain points of validating, monitoring and operating GPU clusters at scale.
  • Design, develop and maintain engineering solutions to solve those pain points systematically.
  • Research in traditional AIOps and the emerging Agentic AI, and leverage it to further reduce the operation toil.
  • Participate in on-call support for systems, platforms built and owned by the team.

Requirements:

  • BS/MS in Computer Science, Engineering, or equivalent experience.
  • 5+ years in software/platform engineering, including 3+ years in ML infrastructure or distributed systems.
  • Experience in software development lifecycle on Linux-based platforms.
  • Strong coding skills in languages such as Python, C++ or Rust.
  • Experience with Docker, Kubernetes, GitLab CI, automated deployments.
  • Experience with AIOps or Agentic AI and apply it successfully in production environment.

Nice to Have:

  • Proficiency with full-stack development: Relational Data Modeling, DB optimization, REST API Semantics, Javascript, CSS, providing API as a service.
  • Passion for building developer-centric platforms with great UX and strong operational reliability.
  • Experience running Slurm or custom scheduling frameworks in production ML environments.
  • Familiarity with GPU computing, Linux systems internals, and performance tuning at scale.