Description
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.
Infrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the core developer tools and on-prem compute platforms used across the entire engineering organization.
As a Senior Infrastructure Reliability Engineer, you’ll own the full lifecycle , patching, upgrades, backups, scaling, and incident response , for services that engineering depends on daily.
Responsibilities
- Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team
- Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray)
- Design and implement automated systems for patching, backups (with validation), and upgrades
- Scale infrastructure to support a fast-growing engineering org
- Use Infrastructure-as-Code (Terraform) to manage environments
- Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
- Define and maintain SLOs for service availability, reliability, and performance
- Build and maintain monitoring, alerting, and observability for developer tool services
- Lead and participate in incident response and root cause analysis
- Work cross-functionally with platform, security, infrastructure (on-prem and cloud), and software teams
Required Qualifications
- Experience operating infrastructure outside of managed cloud services , bare-metal kubernetes and on-prem virtualization (VMware ESXi/vSphere)
- Experience operating production systems using Docker and Kubernetes
- Strong foundational knowledge of Linux (RHEL, Ubuntu)
- Proficiency with at least one cloud platform (AWS, GCP, or Azure)
- Experience managing infrastructure with Infrastructure-as-Code tools (e.g., Terraform/OpenTofu)
- Experience with configuration management tooling (e.g., Ansible, Puppet, Chef)
- Strong problem-solving skills with a focus on automation
- Scripting or software development experience (e.g., Python, Go, Bash)
- Familiarity with CI/CD pipelines and developer tooling
- Ability to own systems end-to-end, from design to incident resolution
- Eligible to obtain and maintain an active U.S. Secret security clearance
Preferred Qualifications
- Experience with RKE2 (or other bare-metal Kubernetes distro such as k3s, kubeadm, or OpenShift) and Cilium
- Prior experience with GitHub Enterprise Server, JFrog Artifactory/Xray, or CircleCI
- Experience with GitOps workflows and tooling (e.g., ArgoCD/FluxCD)
- Experience maintaining highly available, scalable internal tools
- Exposure to security best practices, compliance requirements, or auditing
- Experience supporting large, rapidly scaling engineering organizations
- Experience with monitoring and observability platforms (e.g., Datadog, Prometheus, Grafana)
- Background in SRE or hybrid SWE/DevOps roles
- Experience with on-prem infrastructure operations, reliability, or capacity planning
- Experience operating GPU or HPC infrastructure and workload schedulers (eg., RunAI, Slurm, Kubeflow, Volcano)
Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next.