Description
Thinking Machines Lab is seeking a Network Engineer to own the lowest layers of the network stack for large-scale training and inference. The successful candidate will be responsible for interconnect reliability at scale, across large GPU fabrics.
Responsibilities:
- Reason about and validate GPU network fabric design across deployments.
- Debug RDMA/RoCEv2 across different NIC vendors.
- Own NVLink/NVSwitch interconnect.
- Build host-level network instrumentation and use Linux tooling to build dashboards and alerts.
- Navigate cross-cloud fabric quirks and triage across the NIC, driver, kernel, switch, and workload boundaries.
- Drive escalations with cloud-provider networking teams.
Requirements:
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (Python or Rust).
- Experience operating large-scale clusters and container orchestration systems.
- Comfort operating across the stack and owning projects end-to-end.
- Thrive in a highly collaborative environment.
Preferred qualifications:
- Fluency with host-level debugging tools on Linux.
- Strong communication skills.
- Extensive experience with NVLink/NVSwitch, fabric manager, and IMEX.
- Statistical rigor in reliability reasoning.
- A track record of writing tooling that made the next debugging session meaningfully faster.
Logistics:
- Location: San Francisco, California.
- Compensation: $350,000 - $475,000 USD.
- Visa sponsorship: Available.
- Benefits: Generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/thinkingmachines/jobs/5281215008