Description
Thinking Machines Lab is seeking a Reliability Engineer to ensure the reliability of its GPU supercomputing fleet. The successful candidate will track and resolve hardware issues, diagnose problems, and collaborate with vendors.
Responsibilities
- Investigate, reproduce, and remediate issues across large GPU clusters.
- Own the drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS.
- Automate monitoring of fleet reliability and analyze error rates.
- Drive the firmware lifecycle, including tracking, qualification, and regression analysis.
- Engage vendors directly to resolve issues and manage RMA flows.
- Monitor and improve GPU hardware health signals.
- Write clear postmortems and vendor cases.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (Python or Rust).
- Experience operating large-scale clusters and container orchestration systems.
- Comfort operating across the stack and owning projects end-to-end.
- Ability to thrive in a highly collaborative environment.
- Bias for action with a mindset to take initiative.
Preferred Qualifications
- Fluency with Linux systems and debugging tools.
- Proven statistical rigor in analyzing reliability.
- Track record of debugging problems from application symptoms to root causes in hardware.
- Comfort reading vendor errata, firmware release notes, and kernel changelogs.
- Experience engaging hardware vendors directly.
- Linux kernel literacy.
- Out-of-band management experience.
- Depth in GPU hardware health.
- Significant ownership of the hardware reliability function at scale.
- Strong writing skills.
Logistics
- Location: San Francisco, California.
- Compensation: $350,000 - $475,000 USD per year.
- Visa sponsorship: Available.
- Benefits: Generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/thinkingmachines/jobs/5281223008