# Senior Site Reliability Engineer

**Company**: Duolingo
**Location**: New York, NY
**Experience**: senior
**Job type**: full-time
**Salary**: $182,800-$247,300 USD
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/duolingo/jobs/8784354002?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_aa3672ca-d64

## Description

Our mission at Duolingo is to develop the best education in the world and make it universally available.

As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo's sophisticated distributed systems and products are built and maintained with extraordinary quality, and operated in measurable and scalable ways.

Responsibilities:

- Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence

- Support core infrastructure, understand, diagnose, and debug these systems in production

- Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis

- Maintain and document sustainable postmortem/incident response practices

- Advocate for and implement changes that improve reliability, scalability, and velocity

- Reduce the burden of toil with iterative development of tooling and automation

- Collaborate with engineering teams to release new features and become an authority on our services

Requirements:

- 5+ years of experience within site reliability engineering/DevOps of a product with millions of users

- Experience identifying and solving issues in large-scale distributed systems

- Experience with Java, Kotlin, Python or Go

- An understanding of containerization toolsets and container orchestration technologies (Docker, Mesos, Kubernetes, Nomad, etc)

Exceptional candidates will have:

- Experience in improving automation and tooling to reduce service maintenance toil

- Proven experience driving improvements to incident response processes

- Experience assessing reliability and troubleshooting issues in Dynamo, MySQL, and/or PostgreSQL databases

## Skills

### Required
- Java
- Kotlin
- Python
- Go
- Docker
- Mesos
- Kubernetes
- Nomad
- site reliability engineering
- DevOps

### Nice to have
- Dynamo
- MySQL
- PostgreSQL

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/duolingo/jobs/8784354002?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
