Description
NVIDIA is seeking a Senior Solutions Architect to join the NVIDIA Cloud Partners team, focusing on GPU, NVLink, and infrastructure design. In this role, you will assist with designs and architectures for large-scale GPU-based clusters enabling advanced AI supercomputers and enterprise AI infrastructure.
As a Solutions Architect, you will serve as a key technical expert bridging NVIDIA's GPU and NVLink technology designs with software solutions between engineering and field teams, supporting customers with demanding requirements. You will work on end-to-end cluster design and architecture, performance modeling, validation, and NPI cluster deployments.
Responsibilities:
- Partner with NVIDIA Cloud Partners in GPU cluster design and networking, conveying architecture and optimal process information for building next-generation architectures.
- Guide NVIDIA Cloud Partners in cluster design, balancing design principles and complex limitations for high-performance and supportable GPU clusters.
- Collaborate with NVIDIA Cloud Partners to ensure successful first deployments with new products, including new network architectures and topologies.
- Provide feedback on customer and field perspectives on cluster design and workflows to engineering teams.
- Perform hands-on work to assist NVIDIA Cloud Partners in debugging issues related to cluster design, configuration, and performance.
- Support NPI customer deployments with new GPU and Networking architectures.
Requirements:
- BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, Physics, or a related field (or equivalent experience).
- 8+ years of experience in cluster design, validation, and issue resolution, specifically on GPU and HPC clusters.
- Proven expertise in designing large-scale distributed systems, AI clusters, or HPC infrastructure.
- Ability to translate sophisticated engineering concepts into customer-ready documentation, diagrams, and reference material.
- Expertise in driving customer and partner issues to a close with product and engineering teams.
- Ability to handle multi-functional communications across customer, product team, support team, engineering team, etc.
Preferred Qualifications:
- Experience leading large-scale AI Factory or HPC cluster bring-ups or builds.
- Hands-on experience with NVIDIA products, including GPUs, NVLink, and NVIDIA Networking.
- Knowledge of NCCL, MPI, IMEX, NMX, and collectives in distributed training as it pertains to cluster designs.
- External customer-facing skill-set and background.
- Effective time management and capability to balance multiple tasks and customers while thinking creatively to debug and solve problems.
You will also be eligible for equity and benefits.