Description
NVIDIA is seeking a Senior Cloud Infrastructure and DevOps Solutions Architect to join its Infrastructure Specialist Team. The successful candidate will work on building and advising on large-scale AI and HPC systems, engaging with customers, partners, and cross-functional teams to assess, architect, and guide the implementation of infrastructure projects.
The role involves:
- Owning full-solution validation on partner software stacks, including cluster-wide stability testing and real training-workload acceptance
- Minimizing the time from cluster handover to first production workload
- Ensuring Day 2 production stability at fleet scale, including monitoring, logging, and workload orchestration
- Assessing customer environments and operating heterogeneous open platforms
- Providing consultative guidance and hands-on troubleshooting across the full stack
- Acting as a technical leader for assigned accounts
The ideal candidate will have:
- A BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields
- 8+ years of experience in managing scalable cloud environments and automation engineering roles
- Expertise in cloud, HPC, and GPU technologies, including Kubernetes, AI/ML workloads, Linux, storage systems, automation, and observability
- Strong consultative and communication skills
Nice-to-have skills include:
- Knowledge of CI/CD pipelines and container-based microservices architectures
- Experience with NVIDIA GPU and Network Operators, NVIDIA Base Command Manager, and GPU health and fleet telemetry tooling
- Familiarity with AI-native scheduling and inference frameworks on Kubernetes
- Background with RDMA-based fabrics and DPU/DOCA infrastructure services
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/France-Courbevoie/Senior-Cloud-Infrastructure-and-DevOps-Solutions-Architect_JR2025837