Description
As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC) environment.
You will provide first-line support for HPC users across scheduling, compute, storage, and access-related issues.
Key responsibilities include:
- Troubleshooting job failures, scheduler errors, resource constraints, and performance concerns
- Performing triage of infrastructure incidents and gathering diagnostics
- Monitoring system health, queues, node status, and service availability
- Completing established operational procedures for maintenance, patching, and configuration updates
- Developing and maintaining operational documentation, runbooks, and guidelines
Requirements:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field, or equivalent experience
- 2+ years of experience supporting Linux-based production environments
- Solid Linux systems administration fundamentals (RHEL/CentOS and/or Ubuntu)
- Ability to troubleshoot technical issues methodically
- Experience interacting directly with users in a technical support or operations role
- Strong written communication skills
Preferred qualifications:
- Foundational scripting or automation experience (e.g., Bash or Python)
- Solid understanding of workload schedulers such as LSF, Slurm, or similar systems
- Strong grasp of network computing supporting infrastructure (NFS, automounter, LDAP)
- Experience supporting HPC or large-scale compute environments
- Familiarity with EDA workloads
NVIDIA offers highly competitive salaries and a comprehensive benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/HPC-Operations-Engineer_JR2014178