Description
NVIDIA is seeking a Deep Learning Compiler Engineer - CUDA to join their Architecture group. The successful candidate will work on designing and implementing the DSL and core compiler of a tile-aware GPU programming model for emerging GPU architectures.
Responsibilities:
- Design and implement the DSL and core compiler of a tile-aware GPU programming model for emerging GPU architectures
- Continuously innovate and iterate on the core architecture of the compiler to optimise performance
- Investigate next-generation GPU architectures and provide solutions in the DSL and compiler stack
- Perform performance analysis on emerging AI/LLM workloads and integrate with AI/ML frameworks
Requirements:
- Master's or PhD in Computer Engineering, Computer Science, or related field
- 2+ years of relevant work experience
- Excellent C/C++ programming and software engineering skills
- Good understanding of computer architecture
- Strong problem-solving skills and methodology
- Experience with MLIR/TVM/Triton/LLVM is desirable
- Knowledge of GPU architecture and fast kernel programming skills is a plus
- Familiarity with LLM algorithms or HPC domains is a plus
- Knowledge of multi-GPU distributed communication is a plus
- Excellent oral communication in English is a plus
NVIDIA offers highly competitive salaries and a comprehensive benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Shanghai/Deep-Learning-Compiler-Engineer---CUDA_JR2020800