New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Solutions Architect, Inference Service Providers

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 22 Aug 2026

Description

We are looking for a Senior Solutions Architect to change how the AI inference industry works. As part of NVIDIA's Inference Service Providers (ISP) team, you will support strategic partners who rely on your expertise to push the boundaries of performance on NVIDIA Cloud Partner (NCP) infrastructure.

You will master full-stack inference, covering aspects from smarter routing to rack-scale disaggregation and kernel optimization. Your goal will be to write reference architectures adopted by 10,000+ GPU clusters and educate C-level executives, distinguished engineers, and datacenter developers on scaling their businesses and infrastructure.

Responsibilities:

  • Work with inference partners to build cutting-edge inference services, share the value of NVIDIA's stack, and gather product insights.
  • Develop and operate inference recipes using tools like NVIDIA Dynamo to distribute tasks among GPU workers efficiently.
  • Accelerate inference pipelines with TensorRT-LLM, vLLM, SGLang, and other backends for seamless integration with disaggregated inference.
  • Promote DevOps best practices, such as managing Kubernetes clusters, configuring compute fabrics, and ensuring observability.
  • Provide mentorship and technical leadership to customers and internal teams on deploying disaggregated inference systems and resolving complex issues.

Requirements:

  • 6+ years of experience in Solutions Architecture or similar roles, with 2+ years focused on AI workloads on Kubernetes.
  • Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization and serving.
  • In-depth knowledge of modern inference best practices, including disaggregated serving, multi-tier KV cache management, speculative decoding, quantization, and custom inference kernels.
  • Hands-on experience with full-stack agent design, modern sandboxing, memory and retrieval systems, evaluation, skill design, and governance.
  • Bachelor's degree in Computer Science or Engineering, or equivalent experience.

Preferred Qualifications:

  • Prior experience with NVIDIA inference technologies such as Dynamo, NIXL, and Grove.
  • Familiarity with model customization techniques like SFT, DPO, GRPO, RLVR, and corresponding data curation techniques.
  • Contributions to relevant open-source projects.

NVIDIA offers highly competitive salaries and a comprehensive benefits package, including equity and benefits. You can learn more about what we offer at www.nvidiabenefits.com.