# LSF Scheduler Engineer

**Company**: NVIDIA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $184,000 - $287,500 for Level 4 and $224,000 - $356,500 for Level 5
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-Compute-Platform-Engineer--LSF-_JR2023960?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_588032d9-4ca

## Description

NVIDIA's EDA compute environment relies heavily on its LSF (Load Sharing Facility) setup, which handles millions of cores across federated LSF cells. The company is now consolidating its dual-scheduler estate onto a single LSF platform and seeks an engineer with in-depth knowledge of LSF internals.

As an LSF Scheduler Engineer, you will be responsible for owning scheduler behavior across 15-25 federated LSF cells. This includes:

- Tuning mbatchd and mbschd, analyzing scheduling cycles, and identifying contention patterns as cells approach host-count ceilings

- Diagnosing MultiCluster forwarding problems, such as remote queue sizing and forwarding policy issues

- Designing cell topology and federation as the farm grows, determining what belongs in a cell versus what belongs in a new one

- Collaborating with the IaC engineer to encode scheduler policy into a config schema compatible with MultiCluster

- Working with CAD and methodology teams on workloads that challenge normal assumptions, such as large memory jobs and interactive-versus-batch contention

The ideal candidate will have:

- A BS or MS in Computer Science, Computer Engineering, or equivalent experience

- 8+ years of experience in HPC or large-scale batch compute, with 5+ years on IBM Spectrum LSF

- Demonstrated expertise in LSF internals, including debugging scheduler behavior and explaining scheduling cycles

- Hands-on experience with MultiCluster in a production, multi-site environment

- Strong Linux systems fundamentals, system programming languages, and scripting skills in Python, Perl, and shell

To stand out, you should have:

- Experience working on LSF as a developer or in escalation engineering

- Familiarity with LSF integration points, such as esub, eexec, elim, submit wrappers, RTM, or the LSF APIs

- Background in semiconductor or EDA compute, where license constraints and job constraints intersect

- Experience migrating a production estate off Slurm, PBS, or Grid Engine without a scheduled outage

In terms of compensation, the base salary range is $184,000 - $287,500 for Level 4 and $224,000 - $356,500 for Level 5. You will also be eligible for equity and benefits.

## Skills

### Required
- IBM Spectrum LSF
- MultiCluster
- Linux systems
- Python
- Perl
- shell scripting
- HPC
- large-scale batch compute

### Nice to have
- LSF development
- escalation engineering
- semiconductor
- EDA compute
- Slurm
- PBS
- Grid Engine migration

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-Compute-Platform-Engineer--LSF-_JR2023960?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
