Description
NVIDIA is seeking a Senior HPC Support Engineer - Compute/GPU (DGX Platform) to provide comprehensive solutions for AI hardware and software products.
The successful candidate will be a primary point of contact for customers, assisting with technical questions, debugging, and resolving issues.
Key Responsibilities:
- Resolve sophisticated customer concerns and technical issues related to AI hardware and software products using Linux Operating Systems (Multi-distro)
- Rapidly debug and respond to user-reported issues via telephone, email, or conference calls on the DGX Platform (hardware and software) stack
- Apply industry-standard AI tools to efficiently share debugging results, create internal/external knowledge base articles, and analyze customer issues
- Participate in multi-functional team meetings and provide feedback to engineering and marketing regarding product requirements, customer experience, and support tools
Requirements:
- 5+ years of in-depth customer support and debugging experience for hardware and software products
- Strong organizational skills and ability to prioritize/multi-task easily with limited supervision
- Proven use of established AI technologies in day-to-day job responsibilities
- Established knowledge of Enterprise platform and systems engineering, including Linux triage, servers, and hardware/OS internal issues
- Excellent verbal and written English skills
- Academic degree from an accredited university or college in Networking, Computer Science/Engineering, or Electrical/IT (or equivalent experience)
Preferred Skills:
- Linux System Administration on engineering and networking level, preferably focused on Red Hat Enterprise Linux and Ubuntu distributions
- Deep understanding of at least two of the following: data centers, servers, distributed systems, virtualization, deep learning frameworks, containers/containerization (i.e., Docker, Kubernetes)
- Knowledge and working experience with InfiniBand, RDMA/RoCEv2, and GPU Technology
- Clustering or HPC Data-Center technologies, including Upper Layer Protocols (i.e., MPI, NCCL)
- Shell scripting (Bash/Python)
- Ethernet and Distributed File System Storage technologies
Benefits:
- Highly competitive salaries
- Comprehensive benefits package, including equity and benefits (see www.nvidiabenefits.com/ for more information)
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-HPC-Support-Engineer---Compute-and-GPU-Platform_JR2020113