New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Deep Learning Kernel Software Performance Architect

NVIDIA
Apply →
senior full-time Shanghai

First indexed 9 Jul 2026

Description

NVIDIA is seeking a Deep Learning Kernel Software Performance Architect to optimize GPU kernel performance for state-of-the-art data-center platforms.

You will build automated, data-driven workflows to detect, explain, and prevent performance regressions across key deep learning workloads, partnering closely with kernel developers, compiler teams, infrastructure, and architecture/performance groups.

Responsibilities:

  • Performance analysis, optimization, and debugging:
  • Analyze performance of GPU-accelerated kernels and key deep learning building blocks
  • Identify gaps with baselines or projections and optimize kernel performance to fill gaps
  • Debug performance issues end-to-end
  • Automation and regression infrastructure (Python-heavy):
  • Develop and maintain Python-based automation for performance testing and analysis
  • Design and operate performance test workflows
  • Cross-team collaboration and operating model:
  • Work with kernel developers and compiler teams to ensure performance checks are practical and scalable
  • Collaborate with chip architecture and modeling teams to solidify performance methodology
  • Partner with SWQA and infrastructure teams for execution at scale and reliable pipelines/dashboards

Requirements:

  • Master's or Ph.D. degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or related field
  • Strong programming ability in Python and C/C++ with 2+ years of working experience
  • Solid fundamentals in computer architecture, parallel programming, and performance reasoning
  • Experience with performance analysis workflows
  • Comfortable working across teams and driving issues to decision/closure with clear communication

Nice to Have:

  • Experience with high-performance kernels or math libraries
  • GPU programming and performance experience
  • Strong ML/DL workload understanding
  • Familiarity with simulators or analytical modeling