Description
We are seeking a Senior Performance Software Engineer to join our team focused on developing optimized code to accelerate linear algebra and deep learning operations on NVIDIA GPUs. As a deep learning library performance software engineer, you will be working on delivering high-performance code to NVIDIA's cuDNN, cuBLAS, and TensorRT libraries to accelerate deep learning models.
Responsibilities:
- Writing highly tuned compute kernels to perform core deep learning operations (e.g., matrix multiplies, convolutions, normalizations)
- Following general software engineering best practices, including support for regression testing and CI/CD flows
- Collaborating with teams across NVIDIA:
- CUDA compiler team on generating optimal assembly code
- Deep learning training and inference performance teams on which layers require optimization
- Hardware and architecture teams on the programming model for new deep learning hardware features
Requirements:
- Master's or Ph.D. degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or a related field
- 2+ years of relevant industry experience
- Demonstrated strong C++ programming and software design skills, including debugging, performance analysis, and test design
- Experience with performance-oriented parallel programming, even if it's not on GPUs (e.g., with OpenMP or pthreads)
- Solid understanding of computer architecture and some experience with assembly programming
- Identify bottlenecks, optimize resource utilization, and improve throughput
Preferred qualifications:
- Tuning BLAS or deep learning library kernel code
- CUDA GPU programming
- Numerical methods and linear algebra
- LLVM, TVM tensor expressions, or TensorFlow MLIR
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Shanghai/Senior-Performance-Software-Engineer--Deep-Learning-Libraries_JR2023444