Description
NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance innovative LLM and VLM inference. You will push workloads toward practical performance limits on NVIDIA GPU-accelerated systems.
Your work will span models, serving software, distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver measurable gains in latency, throughput, efficiency, and scale.
Responsibilities:
- Lead end-to-end analysis of LLM/VLM inference processes. Define representative prefill and decode workloads. Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value (KV) cache capacity.
- Build speed-of-light and roofline models to quantify performance headroom. Connect arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to clear optimization hypotheses.
- Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation. Eliminate bottlenecks in host code, CUDA kernels, memory, communication, and scheduling.
- Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism.
- Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement.
- Establish repeatable benchmarks, canonical run records, and performance regression gates.
Requirements:
- More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware.
- Strong programming skills in Python, Rust and/or C++, plus hands-on experience with CUDA or another GPU programming environment.
- Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute.
- Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.
- Practical experience optimizing inference servers and model execution.
- Understanding of distributed systems and networking for accelerated computing.
- BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
Benefits:
- Competitive salaries
- Generous benefits package (www.nvidiabenefits.com)
- Equity and benefits eligibility
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Engineer---AI-Inference-Performance_JR2024262