Description
NVIDIA is seeking a Senior AI Compute Engineer to join its Infrastructure Specialists team. The successful candidate will work on large-scale AI Compute projects, interacting with customers, partners, and internal teams to analyze, define, and implement solutions.
The role involves deploying, managing, and validating AI Compute/HPC infrastructure in Linux-based environments for new and existing customers. The ideal candidate will be a domain expert with customers during planning calls through implementation, provide handover-related documentation, and perform knowledge transfers.
Responsibilities:
- Deploy, manage, and validate AI Compute/HPC infrastructure in Linux-based environments
- Be the domain expert with customers during planning calls through implementation
- Provide handover-related documentation and perform knowledge transfers
- Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements
Requirements:
- 8+ years of experience providing in-depth support and deployment services
- Knowledge and experience with Linux system administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, network-routing/advanced networking
- Cluster management and provisioning technologies for bare-metal servers
- Minimum of a four-year degree from an accredited university or college in Computer Science, Electrical or Computer Engineering or equivalent experience
- Scripting proficiency (Bash, Python, Ansible, etc.)
- Excellent interpersonal skills and the ability to deliver resolutions for customer issues
- Strong organizational skills and ability to prioritize/multi-task easily with limited supervision
- Experience with schedulers such as SLURM, LSF, UGE, etc.
- Ability to travel to customer sites within the United States up to 20% of the time
- Experience with benchmarking tools such as HPL, NCCL tests, MLPerf as well as Kubernetes experience
Preferred qualifications:
- InfiniBand experience
- Experience with GPU focused hardware/software
- Experience with MPI
- Storage technologies such as Lustre or GPFS
- Familiarity with OEM GPU platforms
NVIDIA offers equity and benefits to its employees.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-AI-Compute-Engineer---NVIS_JR2020733