New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Data Infrastructure Engineer, AI Performance

NVIDIA
Apply →
senior full-time Shanghai

First indexed 19 Aug 2026

Description

NVIDIA's AI Computing Architecture team develops analytical models and simulators that guide the design of future GPUs, systems, and AI platforms. We are seeking outstanding software engineers to build and scale the infrastructure behind this simulation ecosystem.

You will develop the distributed execution, data, automation, and visualization platforms that turn architectural models into reliable, reproducible, large-scale studies. You will be sitting at the intersection of distributed systems, performance engineering, data platforms, and architecture.

Responsibilities:

  • Build scalable and reliable infrastructure for running large simulation studies across on-premises compute clusters and cloud environments, improving throughput, resource efficiency, and reproducibility.
  • Establish a unified storage and data platform as the source of truth for simulation configurations, execution state, results, and provenance.
  • Develop self-service analytics and visualization capabilities that help architects explore results and compare performance, power, and design trade-offs.
  • Partner with GPU architects, performance engineers, AI researchers, and software teams to translate emerging AI workloads into reusable simulation platform capabilities.
  • Help define the technical roadmap and engineering practices for NVIDIA's next generation of AI architecture simulation platforms.

Requirements:

  • BS or higher degree in a relevant technical field (CS, EE, CE, Math, etc.).
  • 3+ years of experience building production infrastructure, distributed systems, or data platforms.
  • Strong software engineering and system-design skills in one or more programming languages, with a solid understanding of scalability, reliability, data consistency, and operational trade-offs.
  • Ability to work through ambiguous problems, simplify fragmented systems, and collaborate effectively across engineering and research teams.

Preferred Qualifications:

  • Deep expertise in distributed systems, storage and query engines, compute orchestration, or large-scale data platforms.
  • Experience improving the performance, reliability, or efficiency of compute intensive and data intensive systems.
  • Understanding of LLM inference optimization and the performance characteristics of conversational, agentic, or other emerging AI workloads.
  • Familiarity with architecture simulation, performance modeling, high-performance computing, GPU computing, or AI workloads.
  • A record of strong technical ownership and impact through industry work, research or open-source contributions.