# Site Reliability Engineer

**Company**: Forward
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Salary**: $230,000 - $250,000
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/forwardnetworks/jobs/7800601003?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_456a0a4a-eea

## Description

Forward is transforming how the world's most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry's first network digital twin , a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment.

As a Site Reliability Engineer, you will be building the reliability engineering function at Forward , defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.

## Responsibilities

- Define and drive SRE practices from the ground up , SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use

- Drive the reliability and operational excellence of the Forward SaaS platform

- Build and maintain observability infrastructure , logging, metrics, tracing, and alerting , so the team always knows what's happening before customers do

- Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice

- Partner with engineering teams to embed reliability thinking into the SDLC , capacity planning, load testing, chaos engineering, and production readiness reviews

- Help define and build the SRE team as the company scales , this is a foundational hire with a path to leadership

## Requirements

- 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment

- Proven experience building or significantly maturing an SRE function , not just operating within one someone else built

- Strong fundamentals in networking , TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus

- Hands-on experience with Kubernetes and container orchestration in production environments

- Deep proficiency with observability tooling , Prometheus, Grafana, Datadog, Splunk, or similar

- Strong scripting and automation skills in Python, Bash, or similar

- Experience with cloud platforms , AWS, GCP, or Azure , including infrastructure as code (Terraform, Ansible, or equivalent)

- Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements

- Ability to communicate clearly with both engineering teams and non-technical stakeholders , you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them

## Nice to Have

- Experience supporting enterprise or federal government customers with high availability requirements

- Experience in a foundational or early SRE hire capacity at a growth stage company

## What This Role Is Not

- A pure ops or NOC role , you are building and engineering, not just monitoring

- A siloed function , you will be deeply embedded with product and engineering teams

- A ticket-taker , you will be proactively identifying and solving reliability problems before they become incidents

The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.

## Skills

### Required
- site reliability engineering
- DevOps
- infrastructure engineering
- SaaS
- cloud environment
- networking
- TCP/IP
- DNS
- routing
- switching
- firewalls
- load balancing
- Kubernetes
- container orchestration
- observability tooling
- Prometheus
- Grafana
- Datadog
- Splunk
- scripting
- automation
- Python
- Bash
- cloud platforms
- AWS
- GCP
- Azure
- infrastructure as code
- Terraform
- Ansible

### Nice to have
- Experience supporting enterprise or federal government customers with high availability requirements
- Experience in a foundational or early SRE hire capacity at a growth stage company

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/forwardnetworks/jobs/7800601003?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
