New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Anduril Industries

Site Reliability Engineer

Anduril Industries
Apply →
senior full-time $166,000-$220,000 USD Waltham, Massachusetts

First indexed 27 Jul 2026

Description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.

The Imaging team builds and fields state-of-the-art camera and sensor systems deployed to solve real security challenges for the United States and its allies.

We're looking for a Site Reliability Engineer to join the Imaging team. This is not a product development role, and it isn't a traditional cloud-SRE role either. You are the frontline for keeping fielded imaging systems alive: the person field personnel and customer-support escalations turn to when a deployed system isn't behaving.

Responsibilities

  • Own fielded system reliability. You are responsible for the health and uptime of deployed imaging systems. When issues arise--whether the root cause is in the network, the calibration, an upgrade, or the sensor hardware itself--you triage, diagnose, and drive resolution.
  • Run point on escalations. You'll be first and second line of response for issues coming through our support channels and Anduril's customer-support pipeline (Tier 0 to Mission Success / Product Operations to SRE), acting as the deep-expertise backstop the rest of the funnel escalates to.
  • Turn fires into runbooks. Recurring issues shouldn't be solved live twice. You'll build and maintain runbooks, diagnostics, and self-service tooling that reduce repeat problems and shrink the support load over time.
  • Hold the boundary with engineering. You own everything short of a code fix. When an issue turns out to be a genuine software defect, you'll cleanly reproduce it, document it, and hand it off to the Mission Software Engineers--protecting their focus by keeping frontline support where it belongs.
  • Feed reliability back into the product. You'll turn what you learn in the field into signals that make our systems more supportable--better observability, safer upgrades, more graceful failure.
  • Travel expected approximately 15% of the time for field support and deployment windows.

Requirements

  • 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems--with real ownership after systems ship, not just standing them up
  • Strong Linux fundamentals, including comfort troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained or field environments)
  • Demonstrated ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without always having full visibility into every component
  • Comfortable owning a structured on-call rotation, including scheduled after-hours and weekend coverage
  • Strong written and verbal communication skills, including the ability to run a remote troubleshooting session with a non-technical operator and document what happened
  • Eligibility to obtain and maintain a U.S. Secret clearance

Preferred Qualifications

  • Experience supporting fielded or deployed systems, not just development environments
  • Experience with fielded hardware or sensor systems (EO/IR, optical, or similar), including familiarity with sensor calibration
  • Scripting for diagnostics and automation (Python, Bash, or similar)
  • Familiarity with Nix or NixOS--uncommon, but valuable on our stack
  • Familiarity with systemd service management and observability practices on Linux
  • Familiarity with incident tooling (PagerDuty or equivalent)
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/andurilindustries/jobs/5196517007