# Network Engineer, Supercomputing

**Company**: Thinking Machines Lab
**Location**: San Francisco
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Salary**: $350,000 - $475,000 USD
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/thinkingmachines/jobs/5281215008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_132f4a25-7c8

## Description

Thinking Machines Lab is seeking a Network Engineer to own the lowest layers of the network stack for large-scale training and inference. The successful candidate will be responsible for interconnect reliability at scale, across large GPU fabrics.

Responsibilities:

- Reason about and validate GPU network fabric design across deployments.

- Debug RDMA/RoCEv2 across different NIC vendors.

- Own NVLink/NVSwitch interconnect.

- Build host-level network instrumentation and use Linux tooling to build dashboards and alerts.

- Navigate cross-cloud fabric quirks and triage across the NIC, driver, kernel, switch, and workload boundaries.

- Drive escalations with cloud-provider networking teams.

Requirements:

- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.

- Proficiency in at least one backend language (Python or Rust).

- Experience operating large-scale clusters and container orchestration systems.

- Comfort operating across the stack and owning projects end-to-end.

- Thrive in a highly collaborative environment.

Preferred qualifications:

- Fluency with host-level debugging tools on Linux.

- Strong communication skills.

- Extensive experience with NVLink/NVSwitch, fabric manager, and IMEX.

- Statistical rigor in reliability reasoning.

- A track record of writing tooling that made the next debugging session meaningfully faster.

Logistics:

- Location: San Francisco, California.

- Compensation: $350,000 - $475,000 USD.

- Visa sponsorship: Available.

- Benefits: Generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support.

## Skills

### Required
- Python
- Rust
- Linux
- RDMA
- RoCEv2
- NVLink
- NVSwitch

### Nice to have
- host-level debugging tools
- cloud network primitives
- statistical rigor
- CUDA
- NCCL

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/thinkingmachines/jobs/5281215008?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
