# Manager, Site Reliability Engineering

**Company**: LayerZero
**Location**: Vancouver, BC
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/layerzerolabs/jobs/6176125004?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_1f6d304e-da4

## Description

At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering, dedicated to crafting and maintaining large-scale, resilient systems.

As Manager of SRE, you'll lead a team of engineers responsible for the reliability, performance, and scalability of our blockchain node infrastructure and platform services , while staying technically sharp enough to guide architecture decisions and jump into critical incidents.

## Responsibilities

- Lead and develop a team of SREs , setting technical direction, growth plans, and performance expectations.

- Own the reliability strategy for blockchain node infrastructure across a variety of DLTs, including SLOs, capacity planning, and incident response.

- Partner with Engineering leadership and Product/Platform teams to align reliability investments with business priorities.

- Drive infrastructure-as-code practices, with a focus on Kubernetes and Helm at scale.

- Establish and continuously improve on-call structure, incident detection/triage automation, and postmortem culture.

- Stay hands-on: review designs, dig into complex incidents, and set the technical bar for the team.

## Requirements

- Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experience.

- 6+ years in SRE, DevOps, or infrastructure engineering, including 2+ years directly managing or leading a technical team.

- Deep familiarity with blockchain node infrastructure (validator/full/archive nodes, RPC optimization, etc.).

- Strong proficiency in TypeScript or Golang, with the judgment to know when to write code vs. delegate.

- Advanced knowledge of Unix/Linux internals and distributed systems / high-availability design.

- 3+ years running Kubernetes in production, including Helm chart authoring at scale.

- Track record building or scaling an on-call/incident response process.

- Excellent communication skills , able to represent the team to leadership and hire/retain strong engineers.

## Skills

### Required
- TypeScript
- Golang
- Kubernetes
- Helm
- Unix/Linux
- distributed systems
- high-availability design
- blockchain node infrastructure

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/layerzerolabs/jobs/6176125004?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
