Description
We are seeking a Senior System Software Engineer to work on Dynamo, a Generative AI inference platform. As a Senior System Software Engineer, you will develop open source software to serve inference of trained AI models running on GPUs.
Responsibilities:
- Contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads.
- Innovate in inference-state management for long-running agents, including KV- and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies.
- Build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM.
- Balance a variety of objectives: build robust, scalable, high performance software components, work with team leads to prioritize features and capabilities, load-balance asynchronous requests across available resources, optimize throughput under latency constraints, and integrate the latest open source technology.
Requirements:
- Master's or PhD or equivalent experience.
- 10+ years in Computer Science, Computer Engineering, or related field.
- Ability to work in a fast-paced, agile team environment.
- Excellent Rust/Python programming and software design skills.
- Understanding of modern LLM API semantics.
Preferred qualifications:
- Prior contributions to open-source AI inference frameworks.
- Experience optimizing GPU memory, KV and prefix caches, or high-performance networking for long-context, reasoning, and tool-calling workloads.
- Understanding of LLM-specific inference challenges for agentic workloads.
- Prior experience integrating self-hosted LLM serving stacks with agent harnesses.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-System-Software-Engineer--Agentic-Inference---Dynamo_JR2021665-1