# Senior Software Engineer, AI Resiliency

**Company**: NVIDIA
**Location**: Redmond
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Salary**: $150,000 - $250,000 per year
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-WA-Redmond/Senior-Software-Engineer--AI-Resiliency_JR2018027?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_21448986-082

## Description

We are now looking for a Senior Software Engineer for AI Resiliency! At NVIDIA, we are pushing the boundaries of what's possible in AI. We are currently seeking a Senior Software Engineer to lead the development of AI software resiliency for the most powerful AI supercomputers in the world. As a member of our AI Software Resiliency team, you will play a pivotal role in defining and implementing critical resiliency features for AI supercomputers at a scale of 100,000+ GPUs. Your expertise will be crucial in driving down cluster downtime towards zero, ensuring that our AI systems remain robust and reliable at all times.

**Responsibilities:**

- Develop AI Software Resiliency Features: Implement and optimize software features that improve AI system reliability at a massive scale, such as fast checkpoint-recovery, error detection, error isolation, and straggler/hang detection.

- Hands-On Coding & Optimization: Contribute to large-scale distributed systems with high-quality, production-level C++ and Python code. Enhance performance for AI workloads running on thousands of GPUs.

- Fault Tolerance & Debugging: Work on AI system error handling, implementing techniques to detect silent data corruption (SDC) and other failure scenarios. Assist in developing monitoring tools for proactive failure mitigation.

- Collaborate Across Teams: Work closely with senior engineers, AI researchers, and hardware/software teams to integrate resiliency features into AI frameworks like PyTorch and JAX/XLA.

- Testing & Automation: Develop and implement tests to ensure robustness, scalability, and efficiency of resiliency mechanisms. Contribute to CI/CD pipelines to automate validation of AI workloads.

- Support Production Deployments: Assist in debugging and performance tuning large-scale AI workloads in cloud and HPC environments, ensuring seamless operation of AI training and inference workloads.

**Requirements:**

- You've achieved a Bachelor's, Master's or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.

- Proficiency in C++ and Python, with experience in writing efficient, high-performance code.

- 6+ years of relevant experience

- Strong understanding of distributed systems concepts, parallel programming, and fault tolerance in large-scale computing environments.

- Familiarity with AI frameworks such as PyTorch, JAX/XLA, TensorFlow, or similar.

- Experience with debugging and profiling tools (e.g., gdb, perf, valgrind, NVIDIA Nsight).

- Excellent problem-solving skills and ability to work in a fast-paced, highly collaborative environment.

**Preferred Qualifications:**

- Hands-on experience in training models or working with model training teams.

- Hands-on experience with CUDA, NCCL, or MPI for GPU-accelerated computing, especially at extreme-scale.

- Knowledge of checkpointing strategies, error mitigation, or fault-tolerant computing in AI training.

- Experience working with large-scale AI clusters, HPC environments, or cloud-based AI workloads.

- Strong systems programming skills and experience with low-level performance tuning.

## Skills

### Required
- C++
- Python
- Distributed Systems
- Parallel Programming
- Fault Tolerance
- AI Frameworks
- PyTorch
- JAX/XLA
- TensorFlow
- Debugging Tools
- Profiling Tools

### Nice to have
- CUDA
- NCCL
- MPI
- GPU-Accelerated Computing
- Checkpointing Strategies
- Error Mitigation
- Fault-Tolerant Computing
- Large-Scale AI Clusters
- HPC Environments
- Cloud-Based AI Workloads
- Systems Programming
- Low-Level Performance Tuning

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-WA-Redmond/Senior-Software-Engineer--AI-Resiliency_JR2018027?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
