# Senior Systems Software Engineer, CUDA Driver - Multi-Node and Memory Model

**Company**: NVIDIA
**Location**: Santa Clara
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Systems-Software-Engineer--CUDA-Driver---Multi-Node-and-Memory-Model_JR2014447?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9f2507fc-724

## Description

We're looking for a motivated system software engineer with a deep understanding of device drivers, memory coherency & consistency models, phenomenal C/C++ skills, and an interest in multi-node scalability. As a member of our team, you will use your design abilities, coding expertise, and creativity to deliver the best compute platform in the world. You will craft elegant solutions to exciting problems and shape the future direction of CUDA as you collaborate with your peers across NVIDIA.

**Key Responsibilities:**

- Evangelize, architect, and implement new features related to CUDA's memory model and multi-node scalability geared towards next-gen AI applications and deployments

- Coordinate and drive development efforts across multiple teams

- Help define forward-looking improvements to the CUDA APIs and programming model

- Write effective, maintainable, and well-tested code

- Develop code for multiple operating systems

**Requirements:**

- BS or MS degree in Computer Science, Electrical Engineering or related field (or equivalent experience)

- Strong C and C++ programming skills

- Minimum of 8 years of related development experience

- Experience driving projects across multiple teams

- Experience working with large codebases

- Background with operating system interfaces for threads, process control, and virtual memory

- Experience writing and debugging multithreaded programs

**Nice to Have:**

- Prior experience with parallel computing, PyTorch, low-latency AI inference

- Understanding of system level architecture, such as interconnects, memory hierarchy, interrupts, and memory-mapped IO

- Knowledge of memory coherence and consistency models

- Background with kernel mode development

- Experience with Linux, or Windows Systems Software development

## Skills

### Required
- C++
- CUDA
- device drivers
- memory coherency & consistency models
- multithreaded programs
- operating system interfaces
- process control
- virtual memory

### Nice to have
- parallel computing
- PyTorch
- low-latency AI inference
- system level architecture
- interconnects
- memory hierarchy
- interrupts
- memory-mapped IO
- kernel mode development
- Linux
- Windows Systems Software development

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Systems-Software-Engineer--CUDA-Driver---Multi-Node-and-Memory-Model_JR2014447?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
