Description
NVIDIA is developing processor and system architectures that accelerate deep learning and high-performance computing applications.
We are looking for an expert deep learning system performance architect to join our AI performance modelling, analysis and optimization efforts.
In this position, you will have a chance to work on DL performance modelling, analysis, and optimization on state-of-the-art hardware architectures for various LLM workloads.
Responsibilities:
- Analyze state-of-the-art DL networks (LLM etc.), identify and prototype performance opportunities to influence SW and Architecture team for NVIDIA's current and next-gen inference products
- Develop analytical models for state-of-the-art deep learning networks and algorithms to innovate processor and system architectures design for performance and efficiency
- Specify hardware/software configurations and metrics to analyze performance, power, and accuracy in existing and future uni-processor and multiprocessor configurations
- Collaborate across the company to guide the direction of next-gen deep learning HW/SW by working with architecture, software, and product teams
Requirements:
- BS, MS or PhD in relevant discipline (CS, EE, Math, etc.) or equivalent experience
- 5+ years work experience
- Experience with popular AI models (e.g., LLM and AIGC models)
- Familiarity with typical deep learning SW frameworks (e.g., Torch/JAX/TensorFlow/TensorRT)
- Knowledge and experience on hardware architectures for deep learning applications
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Shanghai/Deep-Learning-Performance-Architect_JR2022306