# Senior Solutions Architect, First Time Deployment Networking InfiniBand - NVIS

**Company**: NVIDIA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Solutions-Architect--First-Time-Deployment-Networking-InfiniBand---NVIS_JR2021241?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_6971d1e4-97a

## Description

NVIDIA is looking for a Senior Solutions Architect to join their team in building large-scale AI/HPC systems. The successful candidate will be responsible for deploying, managing, and validating AI/HPC infrastructure in Linux-based environments for first-of-its-kind NVIDIA hardware.

**Responsibilities:**

- Deploy, manage, and validate AI/HPC infrastructure in Linux-based environments for first-of-its-kind NVIDIA hardware.

- Collaborate with NVIDIA design engineering teams and external customers during planning calls through implementation.

- Create and hand over related documentation and perform knowledge transfers required to support customers.

- Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

**Requirements:**

- 6+ years of experience providing in-depth support and deployment services for hardware and software products.

- Knowledge and experience with Linux system administration, DevOps, process management, package management, task scheduling, kernel management, boot procedures, troubleshooting, performance reporting/optimization/logging, and network-routing/advanced networking.

- Experience in configuring, testing, validating, and issue resolution of LAN and InfiniBand networking.

- Experience with benchmarking tools such as HPL, NCCL tests, MLPERF.

- Scripting proficiency (Bash, Python, Ansible, etc.) and Automation tooling background (Ansible, Puppet, etc.).

- Familiarity with schedulers such as SLURM, LSF, UGE, etc.

- Kubernetes experience.

- Excellent interpersonal communication skills and the ability to deliver resolutions for customer issues.

- A willingness to travel to customer sites within the United States.

- Minimum of a four-year degree from an accredited university or college in Computer Science, Electrical or Computer Engineering, or equivalent experience.

**Nice to Have:**

- Cluster management technologies knowledge.

- Experience with GPU-focused hardware/software.

- Experience with MPI.

- Storage technologies such as Lustre or GPFS.

- Familiarity with Dell and Supermicro GPU platforms.

## Skills

### Required
- Linux system administration
- DevOps
- Networking
- Scripting
- Automation
- Kubernetes
- InfiniBand
- LAN
- HPL
- NCCL tests
- MLPERF
- SLURM
- LSF
- UGE

### Nice to have
- Cluster management technologies
- GPU-focused hardware/software
- MPI
- Lustre
- GPFS
- Dell GPU platforms
- Supermicro GPU platforms

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Solutions-Architect--First-Time-Deployment-Networking-InfiniBand---NVIS_JR2021241?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
