# Reliability Engineer, Supercomputing

**Company**: Thinking Machines Lab
**Location**: San Francisco
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Salary**: $350,000 - $475,000 USD
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/thinkingmachines/jobs/5281223008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_16c9c60a-34f

## Description

Thinking Machines Lab is seeking a Reliability Engineer to ensure the reliability of its GPU supercomputing fleet. The successful candidate will track and resolve hardware issues, diagnose problems, and collaborate with vendors.

## Responsibilities

- Investigate, reproduce, and remediate issues across large GPU clusters.

- Own the drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS.

- Automate monitoring of fleet reliability and analyze error rates.

- Drive the firmware lifecycle, including tracking, qualification, and regression analysis.

- Engage vendors directly to resolve issues and manage RMA flows.

- Monitor and improve GPU hardware health signals.

- Write clear postmortems and vendor cases.

## Requirements

- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.

- Proficiency in at least one backend language (Python or Rust).

- Experience operating large-scale clusters and container orchestration systems.

- Comfort operating across the stack and owning projects end-to-end.

- Ability to thrive in a highly collaborative environment.

- Bias for action with a mindset to take initiative.

## Preferred Qualifications

- Fluency with Linux systems and debugging tools.

- Proven statistical rigor in analyzing reliability.

- Track record of debugging problems from application symptoms to root causes in hardware.

- Comfort reading vendor errata, firmware release notes, and kernel changelogs.

- Experience engaging hardware vendors directly.

- Linux kernel literacy.

- Out-of-band management experience.

- Depth in GPU hardware health.

- Significant ownership of the hardware reliability function at scale.

- Strong writing skills.

## Logistics

- Location: San Francisco, California.

- Compensation: $350,000 - $475,000 USD per year.

- Visa sponsorship: Available.

- Benefits: Generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support.

## Skills

### Required
- Python
- Rust
- large-scale clusters
- container orchestration systems
- Linux systems

### Nice to have
- statistical analysis
- debugging
- vendor engagement
- Linux kernel literacy
- GPU hardware health

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/thinkingmachines/jobs/5281223008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
