# Senior Software Engineer - AI Inference Performance

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer---AI-Inference-Performance_JR2024262?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_46f10db2-b5e

## Description

NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward practical performance limits on NVIDIA GPU-accelerated systems.

Your work will span models, serving software, distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver measurable gains in latency, throughput, efficiency, and scale.

**Responsibilities:**

- Lead end-to-end analysis of LLM/VLM inference processes. Define representative prefill and decode workloads. Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value (KV) cache capacity.

- Build speed-of-light and roofline models to quantify performance headroom. Connect arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to clear optimization hypotheses.

- Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation. Eliminate bottlenecks in host code, CUDA kernels, memory, communication, and scheduling.

- Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism.

- Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement.

- Establish repeatable benchmarks, canonical run records, and performance regression gates.

**Requirements:**

- More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware.

- Strong programming skills in Python, Rust and/or C++, plus hands-on experience with CUDA or another GPU programming environment.

- Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute.

- Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.

- Practical experience optimizing inference servers and model execution.

- Understanding of distributed systems and networking for accelerated computing.

- BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.

**Benefits:**

- Competitive salaries

- Generous benefits package (www.nvidiabenefits.com)

- Equity and benefits eligibility

## Skills

### Required
- Python
- Rust
- C++
- CUDA
- GPU programming
- Speed-of-light analysis
- Roofline models
- Microbenchmarks
- NVIDIA Nsight Systems
- Nsight Compute
- PyTorch Profiler
- Distributed systems
- Networking
- GPU architecture

### Nice to have
- Contributions to high-performance AI projects
- Experience developing AI-agent-supported performance workflows
- Published research
- Conference presentations
- Technical talks
- Blog posts

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer---AI-Inference-Performance_JR2024262?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
