Description
Forward is transforming how the world's most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry's first network digital twin , a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment.
As a Site Reliability Engineer, you will be building the reliability engineering function at Forward , defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.
Responsibilities
- Define and drive SRE practices from the ground up , SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
- Drive the reliability and operational excellence of the Forward SaaS platform
- Build and maintain observability infrastructure , logging, metrics, tracing, and alerting , so the team always knows what's happening before customers do
- Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice
- Partner with engineering teams to embed reliability thinking into the SDLC , capacity planning, load testing, chaos engineering, and production readiness reviews
- Help define and build the SRE team as the company scales , this is a foundational hire with a path to leadership
Requirements
- 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
- Proven experience building or significantly maturing an SRE function , not just operating within one someone else built
- Strong fundamentals in networking , TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
- Hands-on experience with Kubernetes and container orchestration in production environments
- Deep proficiency with observability tooling , Prometheus, Grafana, Datadog, Splunk, or similar
- Strong scripting and automation skills in Python, Bash, or similar
- Experience with cloud platforms , AWS, GCP, or Azure , including infrastructure as code (Terraform, Ansible, or equivalent)
- Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
- Ability to communicate clearly with both engineering teams and non-technical stakeholders , you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them
Nice to Have
- Experience supporting enterprise or federal government customers with high availability requirements
- Experience in a foundational or early SRE hire capacity at a growth stage company
What This Role Is Not
- A pure ops or NOC role , you are building and engineering, not just monitoring
- A siloed function , you will be deeply embedded with product and engineering teams
- A ticket-taker , you will be proactively identifying and solving reliability problems before they become incidents
The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.