# Senior System Architect, Infrastructure Reliability

**Company**: NVIDIA
**Location**: Santa Clara
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-System-Architect--Infrastructure-Reliability_JR2013698?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_8ff16c68-195

## Description

We are seeking a Senior System Architect to join our team and help us solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework that ingests telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time.

Key responsibilities will include:

- Architecting a scalable 'flight recorder' for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.

- Building automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions with system-level events such as OOM kills or NUMA-related hangs.

- Implementing low-overhead tracing mechanisms that provide access to job execution across multi-node Slurm or Kubernetes clusters.

- Developing heuristics and models based on machine learning to classify failures as 'Hardware Fault,' 'Software Bug,' or 'Environment Issue.'

- Working closely with hardware and infrastructure teams to define 'signals of impending failure,' enabling proactive job migration or check-pointing before a crash occurs.

Requirements include:

- A Bachelor's degree in Computer Science or Electrical Engineering (or equivalent experience).

- 6+ years of experience in systems programming.

- Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.

- Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.

- Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.

This is a hybrid position, with eligibility for equity and benefits.

## Skills

### Required
- Distributed Systems
- Machine Learning
- C++
- Python
- x86/ARM Node-Level Metrics
- Cluster Resource Managers
- GPU XID Errors
- PCIe Bus Failures
- CUDA Memory Exceptions

### Nice to have
- Low-Level Diagnostics
- GPU Infrastructure Proficiency
- Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns
- Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-System-Architect--Infrastructure-Reliability_JR2013698?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
