# Senior HPC Support Engineer - Compute and GPU Platform

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: IT
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-HPC-Support-Engineer---Compute-and-GPU-Platform_JR2020113?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_4749c476-298

## Description

NVIDIA is seeking a Senior HPC Support Engineer - Compute/GPU (DGX Platform) to provide comprehensive solutions for AI hardware and software products.

The successful candidate will be a primary point of contact for customers, assisting with technical questions, debugging, and resolving issues.

**Key Responsibilities:**

- Resolve sophisticated customer concerns and technical issues related to AI hardware and software products using Linux Operating Systems (Multi-distro)

- Rapidly debug and respond to user-reported issues via telephone, email, or conference calls on the DGX Platform (hardware and software) stack

- Apply industry-standard AI tools to efficiently share debugging results, create internal/external knowledge base articles, and analyze customer issues

- Participate in multi-functional team meetings and provide feedback to engineering and marketing regarding product requirements, customer experience, and support tools

**Requirements:**

- 5+ years of in-depth customer support and debugging experience for hardware and software products

- Strong organizational skills and ability to prioritize/multi-task easily with limited supervision

- Proven use of established AI technologies in day-to-day job responsibilities

- Established knowledge of Enterprise platform and systems engineering, including Linux triage, servers, and hardware/OS internal issues

- Excellent verbal and written English skills

- Academic degree from an accredited university or college in Networking, Computer Science/Engineering, or Electrical/IT (or equivalent experience)

**Preferred Skills:**

- Linux System Administration on engineering and networking level, preferably focused on Red Hat Enterprise Linux and Ubuntu distributions

- Deep understanding of at least two of the following: data centers, servers, distributed systems, virtualization, deep learning frameworks, containers/containerization (i.e., Docker, Kubernetes)

- Knowledge and working experience with InfiniBand, RDMA/RoCEv2, and GPU Technology

- Clustering or HPC Data-Center technologies, including Upper Layer Protocols (i.e., MPI, NCCL)

- Shell scripting (Bash/Python)

- Ethernet and Distributed File System Storage technologies

**Benefits:**

- Highly competitive salaries

- Comprehensive benefits package, including equity and benefits (see www.nvidiabenefits.com/ for more information)

## Skills

### Required
- Linux System Administration
- AI technologies
- Enterprise platform and systems engineering
- Linux triage
- servers
- Distributed systems
- virtualization
- deep learning frameworks
- containers/containerization

### Nice to have
- InfiniBand
- RDMA/RoCEv2
- GPU Technology
- Clustering or HPC Data-Center technologies
- Shell scripting (Bash/Python)
- Ethernet and Distributed File System Storage technologies

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-HPC-Support-Engineer---Compute-and-GPU-Platform_JR2020113?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
