Description
NVIDIA is seeking a Senior Solution Architect, Networking and Compute Infrastructure to join its Infrastructure Specialist Team. The successful candidate will work on building large-scale AI and HPC systems, interacting with customers, partners, and internal teams to analyze, define, and implement networking projects.
Responsibilities:
- Build AI/HPC infrastructure for new and existing customers
- Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting
- Engage in and improve the whole lifecycle of services,from inception and design through deployment, operation, and refinement
- Develop tooling to automate and manage large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources
- Deploy monitoring solutions for servers, network, and storage
- Perform troubleshooting bottom-up from bare metal, operating system, software stack, and application level
- Develop, redefine, and document standard methodologies to share with customers and internal teams
Requirements:
- BS/MS/PhD or equivalent experience in Computer Science, Data Science, Electrical/Computer Engineering, Physics, Mathematics, or other Engineering fields with at least 5+ years' work or research experience in networking fundamentals, TCP/IP stack, and data centre compute architecture
- Advanced knowledge of HPC, AI, EVPN, BGP, OSPF, VXLAN protocols
- Deep understanding of DC architecture fundamentals such as compute, storage (PFS), and InfiniBand, Ethernet, NVLink
- Experience running HPC performance benchmarks, cluster health checks, and profiling tools to identify infrastructure bottlenecks
- Python programming, bash scripting experience, and advanced Linux knowledge
- Extensive experience delivering automated network provisioning and comfortable with automation and configuration management tools including Jenkins, Ansible, Puppet, Chef, etc.
- Possess solid working knowledge of Ethernet/InfiniBand/RDMA core principles
- Excellent customer-facing and communication skills (verbal and written in both languages)
Nice to Have:
- Knowledge of CPU and/or GPU architecture, including Kubernetes, container-related microservice technologies
- Background with RDMA (InfiniBand or RoCE) fabrics
- Linux or Networking Certifications (e.g., CCNP, CCIE) or NVIDIA-related certifications
- Deep knowledge of observability stack and build experience
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Gurugram/Senior-Solutions-Architect--Networking-and-Compute-Infrastructure_JR2025322