Description
We are seeking an AI Networking Architect to join the Networking Research Group. In this role, you will work at the intersection of AI applications, distributed systems, networking hardware, and software architecture.
You will model the performance of complex AI workloads to identify bottlenecks and recommend system-level optimizations. You will analyze brand-new AI models, distributed training techniques, and inference workloads to understand their infrastructure requirements.
You will build platforms, simulations, and HW platforms, execute AI workloads, and build analytical tools to evaluate trade-offs across compute, memory, storage, and network behavior. You will translate research insights and workload behavior into actionable software, hardware, and networking architecture requirements.
You will partner with architecture, software, and product teams to influence future NVIDIA networking and AI infrastructure roadmaps. You will drive architectural innovation by applying deep workload analysis to real-world advanced machine learning frameworks.
Requirements:
- B.Sc. Or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- 3+ years of relevant industry or research experience.
- Strong machine learning or data science background, with hands-on experience in LLMs, generative AI, or deep learning systems.
- Strong systems-level thinking, capable of estimating end-to-end requirements across the AI stack.
- Shown ability to translate research findings and product requirements into clear software and hardware specifications.
- Excellent research skills, including the ability to digest academic papers, self-learn new domains, and independently test hypotheses.
- Advanced programming skills for performance modeling, data analysis, and prototyping.
- Excellent communication skills, demonstrating proficiency in presenting complex technical findings clearly and confidently.
Nice to Have:
- Experience with distributed training, distributed inference, or large-scale AI serving systems.
- Experience in Agentic programming, and AI tools
- Familiarity with GPU clusters, collective communication, storage systems, or AI networking bottlenecks.