New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA
Apply →
senior full-time

First indexed 19 Sept 2026

Description

NVIDIA is seeking a Senior Cloud Infrastructure and DevOps Solutions Architect to join its Infrastructure Specialist Team. The successful candidate will work on building and advising on large-scale AI and HPC systems, engaging with customers, partners, and cross-functional teams to assess, architect, and guide the implementation of infrastructure projects.

The role involves:

  • Owning full-solution validation on partner software stacks, including cluster-wide stability testing and real training-workload acceptance
  • Minimizing the time from cluster handover to first production workload
  • Ensuring Day 2 production stability at fleet scale, including monitoring, logging, and workload orchestration
  • Assessing customer environments and operating heterogeneous open platforms
  • Providing consultative guidance and hands-on troubleshooting across the full stack
  • Acting as a technical leader for assigned accounts

The ideal candidate will have:

  • A BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields
  • 8+ years of experience in managing scalable cloud environments and automation engineering roles
  • Expertise in cloud, HPC, and GPU technologies, including Kubernetes, AI/ML workloads, Linux, storage systems, automation, and observability
  • Strong consultative and communication skills

Nice-to-have skills include:

  • Knowledge of CI/CD pipelines and container-based microservices architectures
  • Experience with NVIDIA GPU and Network Operators, NVIDIA Base Command Manager, and GPU health and fleet telemetry tooling
  • Familiarity with AI-native scheduling and inference frameworks on Kubernetes
  • Background with RDMA-based fabrics and DPU/DOCA infrastructure services