Description
We're looking for a Senior Performance Compiler Engineer to join our team and work on the open-source Triton compiler project. This opportunity involves working with new technologies and using compilers to improve AI performance on NVIDIA GPUs.
Your work will enable breakthroughs in large language models, agents, and other high-impact AI applications, accelerating both training and inference. You will be immersed in a diverse, supportive environment where everyone is inspired to do their best work, pushing the limits of what's possible.
Responsibilities:
Investigate the latest and future NVIDIA GPU hardware architecture and programming models.
Work on the frontier of AI by understanding advanced algorithms (like attention sinks and MoEs) and numerics (like block-scaled floating point) to identify new opportunities for optimisation.
Design and implement compiler technology using MLIR to optimise high-level kernel descriptions (written in Triton's Python DSL), with a focus on generating efficient, low-level GPU code. When vital, you'll also be able to use inline PTX to hand-tune critical code paths and extract peak performance from the hardware.
Engage in a dynamic, iterative process of optimisation,sometimes starting with the kernel, sometimes with the compiler,to find the most efficient path to peak performance.
Collaborate with teams across NVIDIA, including hardware architects and the CUDA compiler team, to influence future products and ensure we are always operating at maximum efficiency.
Requirements:
Bachelor, Masters or Ph.D. degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or a related field.
8+ years of relevant industry experience in software development.
Demonstrated strong C++ programming and software design skills, with an emphasis on performance analysis and debugging.
Experienced in parallel programming, including CUDA/OpenCL GPU programming or other parallel models such as OpenMP.
Solid understanding of computer architecture and hands-on experience with assembly-level programming.
Preferred Qualifications:
Experience in tuning BLAS or deep learning library kernels.
Background in numerics and linear algebra.
Experience with machine learning compilers like TVM or MLIR.
Contributions to open-source projects, especially in the AI/ML or compiler space.
Familiarity with the latest research in AI algorithms and numerics as well as a strong track record of contributions to open-source projects, particularly in the AI/ML, compiler, or high-performance computing domains.