New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA
Apply →
senior full-time

First indexed 28 Aug 2026

Description

NVIDIA is looking for a Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. The team builds many of the largest and fastest AI/HPC systems in the world.

The role involves working on a dynamic customer-focused team, requiring excellent interpersonal skills. You will interact with customers, partners, and internal teams to analyze, define, and implement large-scale Networking projects.

Responsibilities:

  • Build AI/HPC infrastructure for new and existing customers
  • Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting
  • Engage in and improve the whole lifecycle of services,from inception and design through deployment, operation, and refinement
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health
  • Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements

Requirements:

  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields
  • At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World
  • Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS
  • Solid working knowledge of Ethernet/InfiniBand/RDMA core principles
  • Proficiency in end-to-end IB/Eth cluster deployment, adapter configuration and firmware maintenance, and able to conduct professional performance benchmarking with mainstream RDMA testing tools
  • Capable of independently diagnosing and troubleshooting typical IB/Eth network anomalies
  • Master practical RDMA network optimization strategies such as QP tuning, MTU configuration, and congestion control optimization
  • Hands-on working experience in RDMA-accelerated business scenarios, including distributed storage and high-performance computing clusters
  • Extensive experience delivering automated network provisioning solutions using tools like Ansible, Salt, and Python
  • Ability to develop CI/CD pipelines for network operations
  • Strong written, verbal, and listening skills in English are essential

Benefits:

  • Advanced Linux or Networking Certifications
  • Experience with High-performance computing architectures
  • Understanding of how job schedulers (Slurm, PBS) work
  • Cluster management technologies knowledge (bonus credit for BCM (Base Command Manager))
  • Experience with GPU (Graphics Processing Unit) focused hardware/software