# Senior Systems Software Engineer - Fleet Debuggability

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Systems-Software-Engineer---Fleet-Debuggability_JR2021461?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_31a51660-3b2

## Description

NVIDIA's Datacenter System Software team is seeking a highly motivated Senior Systems Software Engineer to drive Fleet Scale Debuggability end-to-end. You will design, architect, and build infrastructure, tooling, and analytics to collect multi-rack scale logs.

Your responsibilities will include:

- Architecting, designing, and building fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.

- Developing tooling to collect, normalize, and time-align logs from heterogeneous sources, including kernel and driver logs, syslog, Redfish event logs, SEL, firmware, and BMC logs.

- Creating debug and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults.

- Partnering with cross-functional teams, including developers, SWQA, and product engineering, to deliver end-to-end logging solutions.

- Leading the project's open-source release, including keeping internal and public code paths clean, reviewing community contributions, and representing the tooling in upstream discussions.

Requirements:

- 10+ years of experience in the software industry, specializing in system software and/or firmware development.

- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.

- Proven track record of shipping scalable server products or fleet-wide experience.

- Excellent written and oral communication skills, including executive-level reporting.

Preferred qualifications:

- Experience leading debuggability solutions on sophisticated rack-scale compute architectures.

- Familiarity with log and telemetry analytics stacks, such as OpenSearch/ELK, Loki, Prometheus, Grafana, and PagerDuty.

- Hands-on experience with x86/ARM system architecture and coding in C/C++ and Python.

Benefits:

- Equity eligibility

- Comprehensive benefits package

## Skills

### Required
- Python
- RUST
- Linux systems
- kernel and driver logs
- syslog
- Redfish event logs
- SEL
- firmware and BMC logs
- log parsing
- normalization
- structured logging

### Nice to have
- Experience leading debuggability solutions on sophisticated rack-scale compute architectures
- Familiarity with log and telemetry analytics stacks
- Hands-on experience with x86/ARM system architecture and coding in C/C++ and Python

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Systems-Software-Engineer---Fleet-Debuggability_JR2021461?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
