New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Software Engineer - AI Inference Performance

NVIDIA
Apply →
senior full-time Santa Clara, CA

First indexed 27 Aug 2026

Description

NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward practical performance limits on NVIDIA GPU-accelerated systems.

Your work will span models, serving software, distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver measurable gains in latency, throughput, efficiency, and scale.

Responsibilities:

  • Lead end-to-end analysis of LLM/VLM inference processes. Define representative prefill and decode workloads. Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value (KV) cache capacity.
  • Build speed-of-light and roofline models to quantify performance headroom. Connect arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to clear optimization hypotheses.
  • Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation. Eliminate bottlenecks in host code, CUDA kernels, memory, communication, and scheduling.
  • Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism.
  • Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement.
  • Establish repeatable benchmarks, canonical run records, and performance regression gates.

Requirements:

  • More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware.
  • Strong programming skills in Python, Rust and/or C++, plus hands-on experience with CUDA or another GPU programming environment.
  • Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute.
  • Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.
  • Practical experience optimizing inference servers and model execution.
  • Understanding of distributed systems and networking for accelerated computing.
  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.

Benefits:

  • Competitive salaries
  • Generous benefits package (www.nvidiabenefits.com)
  • Equity and benefits eligibility