Description
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.
The Imaging team builds and fields state-of-the-art camera and sensor systems deployed to solve real security challenges for the United States and its allies.
We're looking for a Site Reliability Engineer to join the Imaging team. This is not a product development role, and it isn't a traditional cloud-SRE role either. You are the frontline for keeping fielded imaging systems alive: the person field personnel and customer-support escalations turn to when a deployed system isn't behaving.
Responsibilities
- Own fielded system reliability. You are responsible for the health and uptime of deployed imaging systems. When issues arise--whether the root cause is in the network, the calibration, an upgrade, or the sensor hardware itself--you triage, diagnose, and drive resolution.
- Run point on escalations. You'll be first and second line of response for issues coming through our support channels and Anduril's customer-support pipeline (Tier 0 to Mission Success / Product Operations to SRE), acting as the deep-expertise backstop the rest of the funnel escalates to.
- Turn fires into runbooks. Recurring issues shouldn't be solved live twice. You'll build and maintain runbooks, diagnostics, and self-service tooling that reduce repeat problems and shrink the support load over time.
- Hold the boundary with engineering. You own everything short of a code fix. When an issue turns out to be a genuine software defect, you'll cleanly reproduce it, document it, and hand it off to the Mission Software Engineers--protecting their focus by keeping frontline support where it belongs.
- Feed reliability back into the product. You'll turn what you learn in the field into signals that make our systems more supportable--better observability, safer upgrades, more graceful failure.
- Travel expected approximately 15% of the time for field support and deployment windows.
Requirements
- 3+ years in SRE, DevOps, field/systems engineering, or production support of deployed hardware/software systems--with real ownership after systems ship, not just standing them up
- Strong Linux fundamentals, including comfort troubleshooting real networking issues (IP, routing, VPNs, connectivity in constrained or field environments)
- Demonstrated ability to diagnose and resolve issues across system boundaries (networking, services, hardware interaction) without always having full visibility into every component
- Comfortable owning a structured on-call rotation, including scheduled after-hours and weekend coverage
- Strong written and verbal communication skills, including the ability to run a remote troubleshooting session with a non-technical operator and document what happened
- Eligibility to obtain and maintain a U.S. Secret clearance
Preferred Qualifications
- Experience supporting fielded or deployed systems, not just development environments
- Experience with fielded hardware or sensor systems (EO/IR, optical, or similar), including familiarity with sensor calibration
- Scripting for diagnostics and automation (Python, Bash, or similar)
- Familiarity with Nix or NixOS--uncommon, but valuable on our stack
- Familiarity with systemd service management and observability practices on Linux
- Familiarity with incident tooling (PagerDuty or equivalent)
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/andurilindustries/jobs/5196517007