# Senior Site Reliability Engineer

**Company**: Formation Bio
**Location**: New York, NY
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $185,500 - $232,000
**Category**: Engineering
**Industry**: Healthcare

**Apply**: https://job-boards.greenhouse.io/formationbio/jobs/8213744?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_e5179f64-985

## Description

Formation Bio is seeking a Senior Site Reliability Engineer to build and operate the infrastructure, delivery systems, and operational practices that allow their engineering organization to ship reliable software quickly and safely.

The ideal candidate will have 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline. They will work across cloud infrastructure, developer platforms, observability, and production workloads, including product applications, internal tools, data systems, and ML and AI workloads.

Responsibilities:

- Own the infrastructure and operational platform for shared engineering workloads

- Build and operate secure, observable, reliable infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software

- Research, develop, and maintain core AWS infrastructure and additional cloud outposts

- Create, review, maintain, and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns

- Establish strong operational practices, including SLOs, monitoring, alerting, runbooks, incident response, advanced diagnostics, root cause analysis, and post-incident follow-through

- Work with Product engineering, Data Engineering, and Data Science to evaluate, negotiate, and implement architecture and infrastructure for product software, data systems, model training, and inference

- Use AI tools to accelerate infrastructure development, investigate incidents, improve documentation, build automation, and make operational improvements

- Participate in the support rotation and incident response

- Write and review requirements, design documents, and operating procedures

Requirements:

- 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline

- Production experience operating cloud infrastructure and distributed systems

- Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation

- Experience with AWS and Snowflake

- Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking

- Daily fluency with AI tools, including LLMs and agentic coding systems

- Exceptional collaboration and communication skills

## Skills

### Required
- Site Reliability Engineering
- cloud infrastructure
- DevOps
- AWS
- Snowflake
- Docker
- GitHub
- Kubernetes
- Python
- Terraform
- OpenTofu
- AI tools

### Nice to have
- Azure
- GCP
- Vercel
- Terragrunt
- MLOps infrastructure
- workflow orchestration
- model serving

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/formationbio/jobs/8213744?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
