# Senior Solutions Architect, First Time Deployment Validation - NVIS

**Company**: NVIDIA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Solutions-Architect--First-Time-Deployment-Validation---NVIS_JR2021226?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_05a21639-783

## Description

We are seeking an ambitious Senior Solutions Architect to drive validation of NVIDIA AI factories from first rack power-on through customer handoff. You will be embedded in launches from the start, running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters using NCCL and collectives (AllReduce, AllToAll) to validate performance and scalability.

**Key Responsibilities:**

- Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.

- Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.

- Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.

- Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.

- Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.

- Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks.

- Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.

- Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.

- Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.

- Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.

**Requirements:**

- Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.

- More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.

- Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.

- Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.

- Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.

- Proficiency with Python and Shell/Bash for scripting, automation, and tooling.

- Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).

- Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.

- Strong communication skills and the ability to work effectively with cross-functional teams.

**Nice to Have:**

- Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.

- Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.

- Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.

- Experience building automation and CI-style pipelines for running and validating benchmarks at scale.

- Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.

## Skills

### Required
- Linux
- HPC
- AI/ML
- NCCL
- Python
- Shell/Bash
- benchmarking
- observability

### Nice to have
- AI factory
- large-scale AI infrastructure
- HPC performance engineering
- SRE
- systems performance analysis
- observability stacks
- automation
- CI-style pipelines

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Solutions-Architect--First-Time-Deployment-Validation---NVIS_JR2021226?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
