Description
NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem.
Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time.
You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.
Our job is to work hand in hand with NCPs to solve real problems and drive real optimizations, prove the answer, and turn it into something the next partner can use!
Responsibilities:
- Solve hard Day 2 operations problems at scale. Work alongside partner engineers to find the cause, prototype an approach, validate it under representative load, and leave behind a practice their team can operate.
- Make new technology Day 2 ready. Help partners prepare the operating model for new NVIDIA platforms, capacity, services, and use cases before customers depend on them, and help drive adoption in live environments without degrading service.
- Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per token to show where the cloud is losing performance or margin - and whether the fix worked.
- Raise each partner's Day 2 maturity. Identify and help close the gaps that matter across people, process, tooling, telemetry, security, and incident response.
- Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows that other NCPs can integrate into their standard operating model.
- Create the feedback loop only NVIDIA can. Spot patterns across partners early and bring clear field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level.
Requirements:
- BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.
- 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
- Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.
- Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure.
- Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.
- Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.
Benefits:
- Competitive salaries
- Generous benefits package
- Equity
How to Apply:
Applications for this job will be accepted at least until August 25, 2026.