# Technical Support Engineer - Slurm

**Company**: NVIDIA
**Experience**: senior
**Job type**: full-time
**Salary**: 108,000 USD - 172,500 USD for Level 3, and 120,000 USD - 207,000 USD for Level 4
**Category**: IT
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-TX-Austin/Senior-Technical-Support-Engineer---Slurm_JR2025573?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_09096eb4-2c8

## Description

NVIDIA is seeking a Technical Support Engineer to support Slurm for its customers. The successful candidate will join a dedicated team of Slurm subject-matter experts, owning sophisticated support cases and helping customers run reliable, efficient, and highly scalable systems.

**Job Description:**

As a Technical Support Engineer, you will:

- Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters.

- Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.

- Solve Slurm configuration and policy features, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.

- Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required.

- Isolate problems across Slurm and its surrounding dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.

- Advise customers on Slurm configuration, upgrades, operational practices, managing system resources, and safe recovery from production incidents.

- Collaborate with engineering teams by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.

- Develop guides, knowledge-base articles, diagnostic tools, and internal training that strengthen Slurm expertise across the support organization.

**Requirements:**

- BS degree in Computer Science, Engineering, or a related field, or equivalent experience.

- 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments including business-critical outage incidents.

- Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.

- Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution.

- In-depth Linux system-administration and solve experience, including systemd, cgroups, authentication, networking, and database-backed services.

- Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources.

- Strong analytical and research skills, showing proficiency in distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems.

- Excellent written and verbal communication skills, including the ability to turn detailed technical findings into clear explanations and actionable recommendations.

**Nice to Have:**

- Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs.

- Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues.

- Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation.

- Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity.

- Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform.

**What We Offer:**

- Highly competitive salaries and a comprehensive benefits package.

- Base salary range: 108,000 USD - 172,500 USD for Level 3, and 120,000 USD - 207,000 USD for Level 4.

- Equity and benefits.

## Skills

### Required
- Slurm
- Linux
- HPC
- AI
- GPU
- cluster management

### Nice to have
- containers
- HPC integration technologies
- Slurm source code
- plugins
- SPANK
- Lua job-submit plugins

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-TX-Austin/Senior-Technical-Support-Engineer---Slurm_JR2025573?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
