New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Technical Lead - Network System Validation

NVIDIA
Apply →
senior full-time

First indexed 16 Sept 2026

Description

NVIDIA is seeking a Technical Lead to join the Network System Validation group. The successful candidate will lead the validation of advanced networking solutions across complex AI cluster environments. This is a hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence.

Key Responsibilities:

  • Review system and product requirements, design validation methodologies, and develop comprehensive test plans for networking technologies in large-scale AI cluster solutions
  • Develop and maintain benchmarks, automation tools, and scripts for test execution, environment setup, log collection, and data analysis
  • Lead end-to-end investigation of complex issues, reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks
  • Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
  • Collaborate with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components
  • Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
  • Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes

Requirements:

  • B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience
  • 12+ years of experience in networking, system validation, or related domains
  • Proven experience debugging complex production systems
  • Ability to read, debug, and reason about C/C++ code
  • Strong scripting and automation experience using Python, Bash, and/or Ansible
  • Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress
  • Ability to drive technical alignment across teams and make high-quality architectural decisions at speed

Preferred Qualifications:

  • Experience with large-scale clusters or distributed systems
  • Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)
  • Background in performance analysis, Kubernetes, or cloud environments
  • Background in chaos testing, fault injection, or simulation systems