Description
NVIDIA is seeking a Senior Solutions Architect to partner with account and OEM partner teams to qualify opportunities, design solutions, run technical evaluations, and prove value to customers building large-scale AI and HPC platforms.
What you will be doing
- Work closely with account managers to understand customer requirements, position SONiC + Spectrum whitebox switch solutions, and shape technical win strategies for AI and HPC opportunities.
- Design end-to-end architectures for customer proposals, including GPU server configurations, storage connectivity, and SONiC-based leaf-spine fabrics with RDMA/RoCE.
- Build and validate hands-on demos and POCs: deploy SONiC switches and GPU servers, configure networking, install software stacks, and run benchmarks to prove performance, scalability, and reliability.
- Provide support for large-scale production SONiC switch clusters.
- Own and manage customized SONiC software projects on Spectrum white-box switches.
- Collaborate with internal engineering, product, and OEM partners to resolve complex issues in firmware, drivers, OS, routing, and GPU/network performance, then bring fixes back to active POCs.
What we need to see
- BS/BA in Computer Science, Electrical/Computer Engineering, or equivalent practical experience.
- 6+ years in data center or cloud infrastructure roles (solutions architect, systems engineer, network engineer) with direct exposure to presales or customer-facing technical work.
- Solid background in data center networking for AI workloads: leaf-spine designs, high-bandwidth/low-latency fabrics, RDMA/RoCE, and ideally InfiniBand; comfortable configuring and debugging these in lab and customer environments.
- Practical knowledge of SONiC: installing and upgrading, configuring interfaces and routing (BGP/EVPN/VXLAN), using monitoring/telemetry tools, and troubleshooting real incidents.
- Strong, practical understanding of GPU server architecture: CPU/GPU balance, memory bandwidth, PCIe/NVLink topology, storage and NIC placement, and power/cooling at rack level.
- Hands-on experience designing, deploying, or operating AI/HPC clusters using GPU-accelerated servers (on-prem or cloud), including real involvement in sizing, configuration, and performance tuning.
- Excellent communication and presentation skills: able to explain complex technical topics to both highly technical engineers and non-technical decision makers, and write clear design documents and POC reports. Fluent written and spoken English is required.
Ways to stand out from the crowd
- Contributions to open-source SONiC, networking, or AI infrastructure projects, or published talks/whitepapers on AI data center design and performance tuning.
- Coding experience on SONiC.
- Coding experience with NCCL, NIXL or other collective communication library, RDMA applications, or performance-critical distributed training frameworks, giving you credibility when discussing low-level performance with customer engineers.
- Hands-on deployments of SONiC in production or large lab environments, especially integrated with GPU clusters and RDMA/RoCE fabrics.
- Hands-on experience with NVIDIA networking products and solutions (Spectrum-X, InfiniBand, Cumulus Linux, etc.).
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/Senior-Solutions-Architect--Networking---GPU-System_JR2022712