Description
NVIDIA is seeking a system software engineer to develop low-level diagnostic software for next-generation data center GPUs and rack-scale AI systems.
The successful candidate will own well-scoped components of the diagnostic software from design through implementation, validation, productization, and field support. They will collaborate with hardware architects, driver developers, silicon-validation engineers, manufacturing teams, and field engineers to bring up new hardware and diagnose difficult system failures.
Key Responsibilities:
- Develop diagnostic and stress software in C/C++ and Python for complex hardware systems.
- Collaborate with hardware blocks, firmware, Linux device drivers, registers, telemetry, and low-level debugging tools.
- Bring up and validate new silicon and system features using pre-production hardware and software.
- Create targeted tests for compute engines, memory and cache subsystems, DMA engines, PCIe/NVLink interfaces, power, and thermal behavior.
- Investigate hardware and software failures involving memory errors, ECC, data integrity, performance, thermals, voltage/frequency behavior, and high-speed interfaces.
- Contribute to diagnostic and stress workloads ranging from low-level tests for GPU hardware to higher-level AI workloads.
- Use modern development and analysis tools, including AI-assisted tools, to accelerate coding, debugging, test creation, and failure analysis.
Requirements:
- BS or MS degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent experience.
- 5+ years of experience in embedded software, firmware, Linux device drivers, systems software, hardware validation, diagnostics, or silicon bring-up.
- Strong programming skills in C and C++, plus working proficiency in Python.
- Experience developing software that interacts with hardware, firmware, device drivers, hardware registers, or low-level interfaces.
- Experience creating diagnostics, validation tests, stress tests, manufacturing tests, or other software used to isolate hardware or system failures.
- Understanding of fundamental computer architecture concepts such as memory, caches, interrupts, DMA, buses, and device I/O.
- Strong debugging and problem-solving skills, including the ability to investigate failures across hardware and software boundaries.
- Ability to take ownership of a well-scoped problem and drive it to completion while collaborating with a technical lead and multi-functional teams.
- Good written and verbal communication skills.
NVIDIA offers highly competitive salaries and a comprehensive benefits package, including equity and benefits.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-NC-Durham/System-Software-Engineer---Data-Center-Compute-Diagnostics_JR2022612