New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior HPC Support Engineer - Compute and GPU Platform

NVIDIA
Apply →
remote senior full-time

First indexed 25 Jun 2026

Description

NVIDIA is seeking a Senior HPC Support Engineer - Compute/GPU (DGX Platform) to provide comprehensive solutions for AI hardware and software products.

The successful candidate will be a primary point of contact for customers, assisting with technical questions, debugging, and resolving issues.

Key Responsibilities:

  • Resolve sophisticated customer concerns and technical issues related to AI hardware and software products using Linux Operating Systems (Multi-distro)
  • Rapidly debug and respond to user-reported issues via telephone, email, or conference calls on the DGX Platform (hardware and software) stack
  • Apply industry-standard AI tools to efficiently share debugging results, create internal/external knowledge base articles, and analyze customer issues
  • Participate in multi-functional team meetings and provide feedback to engineering and marketing regarding product requirements, customer experience, and support tools

Requirements:

  • 5+ years of in-depth customer support and debugging experience for hardware and software products
  • Strong organizational skills and ability to prioritize/multi-task easily with limited supervision
  • Proven use of established AI technologies in day-to-day job responsibilities
  • Established knowledge of Enterprise platform and systems engineering, including Linux triage, servers, and hardware/OS internal issues
  • Excellent verbal and written English skills
  • Academic degree from an accredited university or college in Networking, Computer Science/Engineering, or Electrical/IT (or equivalent experience)

Preferred Skills:

  • Linux System Administration on engineering and networking level, preferably focused on Red Hat Enterprise Linux and Ubuntu distributions
  • Deep understanding of at least two of the following: data centers, servers, distributed systems, virtualization, deep learning frameworks, containers/containerization (i.e., Docker, Kubernetes)
  • Knowledge and working experience with InfiniBand, RDMA/RoCEv2, and GPU Technology
  • Clustering or HPC Data-Center technologies, including Upper Layer Protocols (i.e., MPI, NCCL)
  • Shell scripting (Bash/Python)
  • Ethernet and Distributed File System Storage technologies

Benefits:

  • Highly competitive salaries
  • Comprehensive benefits package, including equity and benefits (see www.nvidiabenefits.com/ for more information)