New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA
Apply →
remote senior full-time

First indexed 27 Jun 2026

Description

We are seeking a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW. You will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams and CSP engineering teams to ensure reliable deployment, monitoring, and operation of these systems at fleet scale.

Key Responsibilities:

  • Drive rack-scale SW/FW architecture alignment across CSP engagements, including fabric management software, link health monitoring, and error handling and recovery
  • Drive technical work streams with CSP engineering teams on rack-scale system software, ensuring they understand fabric management, NVSwitch behavior, and error handling and recovery policies
  • Capture and synthesize CSP engineering feedback on rack-scale system software and champion that feedback into NVIDIA's architecture decisions
  • Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development
  • Identify cross-CSP patterns in rack-scale SW/FW issues and drive documentation, tooling, and test strategy improvements

Requirements:

  • 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering
  • Deep understanding of rack-scale system software challenges, including multi-component coordination, error propagation, and health monitoring
  • Experience with fabric management software, cluster management, or system-level orchestration frameworks
  • Understanding of error handling and recovery design patterns in distributed systems
  • Experience with health monitoring and telemetry systems

Nice to Have:

  • Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software
  • Background in system software for large-scale clusters at a hyperscaler
  • Experience crafting error handling and recovery frameworks for multi-component systems
  • Familiarity with GPU or accelerator fleet operations

What We Offer:

  • Competitive salary
  • Equity
  • Benefits