# Principal Software Engineer, Rack-Scale System Software — CSP Engagements

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--Rack-Scale-System-Software---CSP-Engagements_JR2020316?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_3475e81b-29b

## Description

We are seeking a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW. You will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams and CSP engineering teams to ensure reliable deployment, monitoring, and operation of these systems at fleet scale.

**Key Responsibilities:**

- Drive rack-scale SW/FW architecture alignment across CSP engagements, including fabric management software, link health monitoring, and error handling and recovery

- Drive technical work streams with CSP engineering teams on rack-scale system software, ensuring they understand fabric management, NVSwitch behavior, and error handling and recovery policies

- Capture and synthesize CSP engineering feedback on rack-scale system software and champion that feedback into NVIDIA's architecture decisions

- Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development

- Identify cross-CSP patterns in rack-scale SW/FW issues and drive documentation, tooling, and test strategy improvements

**Requirements:**

- 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering

- Deep understanding of rack-scale system software challenges, including multi-component coordination, error propagation, and health monitoring

- Experience with fabric management software, cluster management, or system-level orchestration frameworks

- Understanding of error handling and recovery design patterns in distributed systems

- Experience with health monitoring and telemetry systems

**Nice to Have:**

- Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software

- Background in system software for large-scale clusters at a hyperscaler

- Experience crafting error handling and recovery frameworks for multi-component systems

- Familiarity with GPU or accelerator fleet operations

**What We Offer:**

- Competitive salary

- Equity

- Benefits

## Skills

### Required
- system software engineering
- distributed systems
- fabric management software
- error handling and recovery
- health monitoring and telemetry

### Nice to have
- NVIDIA NVSwitch
- NVOS
- GPU fabric management software
- large-scale clusters
- GPU or accelerator fleet operations

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--Rack-Scale-System-Software---CSP-Engagements_JR2020316?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
