# Site Reliability Engineer

**Company**: Anduril Industries
**Location**: Waltham, Massachusetts
**Experience**: senior
**Job type**: full-time
**Salary**: $166,000-$220,000 USD
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/andurilindustries/jobs/5196517007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_5ddb82c0-154

## Description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.

The Imaging team builds and fields state-of-the-art camera and sensor systems deployed to solve real security challenges for the United States and its allies.

We're looking for a Site Reliability Engineer to join the Imaging team. This is not a product development role, and it isn't a traditional cloud-SRE role either. You are the frontline for keeping fielded imaging systems alive: the person field personnel and customer-support escalations turn to when a deployed system isn't behaving.

## Responsibilities

- Own fielded system reliability. You are responsible for the health and uptime of deployed imaging systems. When issues arise--whether the root cause is in the network, the calibration, an upgrade, or the sensor hardware itself--you triage, diagnose, and drive resolution.

- Run point on escalations. You'll be first and second line of response for issues coming through our support channels and Anduril's customer-support pipeline (Tier 0 to Mission Success / Product Operations to SRE), acting as the deep-expertise backstop the rest of the funnel escalates to.

- Turn fires into runbooks. Recurring issues shouldn't be solved live twice. You'll build and maintain runbooks, diagnostics, and self-service tooling that reduce repeat problems and shrink the support load over time.

- Hold the boundary with engineering. You own everything short of a code fix. When an issue turns out to be a genuine software defect, you'll cleanly reproduce it, document it, and hand it off to the Mission Software Engineers--protecting their focus by keeping frontline support where it belongs.

- Feed reliability back into the product. You'll turn what you learn in the field into signals that make our systems more supportable--better observability, safer upgrades, more graceful failure.

- Travel expected approximately 15% of the time for field support and deployment windows.

## Requirements

- 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems--with real ownership after systems ship, not just standing them up

- Strong Linux fundamentals, including comfort troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained or field environments)

- Demonstrated ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without always having full visibility into every component

- Comfortable owning a structured on-call rotation, including scheduled after-hours and weekend coverage

- Strong written and verbal communication skills, including the ability to run a remote troubleshooting session with a non-technical operator and document what happened

- Eligibility to obtain and maintain a U.S. Secret clearance

## Preferred Qualifications

- Experience supporting fielded or deployed systems, not just development environments

- Experience with fielded hardware or sensor systems (EO/IR, optical, or similar), including familiarity with sensor calibration

- Scripting for diagnostics and automation (Python, Bash, or similar)

- Familiarity with Nix or NixOS--uncommon, but valuable on our stack

- Familiarity with systemd service management and observability practices on Linux

- Familiarity with incident tooling (PagerDuty or equivalent)

## Skills

### Required
- Linux
- networking
- troubleshooting
- communication
- documentation
- U.S. Secret clearance

### Nice to have
- fielded systems
- sensor systems
- scripting
- Nix
- NixOS
- systemd
- incident tooling

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/andurilindustries/jobs/5196517007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
