New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Solutions Architect, Generative AI Deployment and AIOps

NVIDIA
Apply →
remote senior full-time Santa Clara, CA

First indexed 29 Aug 2026

Description

NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers that are building solutions with our newest AI technology.

As a Senior Solutions Architect, Generative AI Deployment and AIOps, you will become a trusted technical advisor with our customers and work on exciting projects and proof-of-concepts focused on inference for Generative AI and Large Language Models (LLMs).

Responsibilities:

  • Partner with solution architects, engineering, product, and business teams to understand their strategies and technical needs and help define high-value solutions
  • Engage with developers, scientific researchers, and data scientists to gain experience across a range of technical areas
  • Collaborate with lighthouse customers and industry-specific solution partners targeting our computing platform
  • Work closely with customers to help them adopt and build creative solutions using NVIDIA technology and MLOps solutions
  • Analyze performance and power efficiency of AI inference workloads on Kubernetes
  • Travel to conferences and customers may be required (20%)

Requirements:

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
  • 8+ years of hands-on experience with Deep Learning frameworks such as PyTorch and TensorFlow
  • Strong fundamentals in programming, optimizations, and software design, especially in Python
  • Proficiency in problem-solving and debugging skills in GPU orchestration and Multi-Instance GPU (MIG) management within Kubernetes environments
  • Experience with containerization and orchestration technologies, monitoring, and observability solutions for AI deployments
  • Excellent knowledge of the theory and practice of LLM and DL inference
  • Excellent presentation, communication, and collaboration skills

Preferred Qualifications:

  • Prior experience with DL training at scale, deploying or optimizing DL inference in production
  • Experience with NVIDIA GPUs and software libraries such as NVIDIA NIM, Dynamo, TensorRT, TensorRT-LLM
  • Excellent C/C++ programming skills, including debugging, profiling, code optimization, performance analysis, and test design
  • Familiarity with parallel programming and distributed computing platforms