Description
About the Role
As a Staff Site Reliability Engineer at EarnIn, you will guide technical direction for reliability across critical services, relying on AI-assisted workflows as key tools to reduce toil, speed incident response, improve production readiness, and enhance the operational quality of the engineering organization.
Responsibilities
- Act as a Staff-level technical leader: define standards, architect solutions, mentor engineers, influence cross-team efforts, and construct reusable systems and practices that multiply your impact.
- Embed AI-first thinking into reliability practices, leveraging AI to streamline alert triage, accelerate incident investigation, automate runbooks, retrieve operational knowledge, enhance postmortem quality, track corrective actions, quantify reliability with scorecards, detect capacity risks, and analyze architectural risks.
- Maintain human ownership and engineering judgment at the center of operations. AI aids engineers by speeding context gathering, clarifying reasoning, and reducing repetition, but it does not replace accountability.
- Collaborate with SRE, product engineering, infrastructure, security, and leadership teams to embed reliability, making it easy to adopt and impossible to ignore.
Key Areas of Ownership
Reliability Strategy and Standards
- Define and evolve reliability standards across critical services, including SLIs, SLOs, error budgets, production readiness, observability, incident response, and resilience patterns.
- Establish a reliability operating model that clarifies service ownership, operational expectations, and decision-making around reliability tradeoffs for product engineering teams.
- Use AI-assisted analysis to interpret reliability trends, detect weak operational signals, highlight capacity risks using pattern recognition, and generate actionable reliability scorecards for teams.
AI-first Incident Response and Operational Workflows
- Overhaul key stages of the incident lifecycle to achieve faster detection, sharper triage, richer context retrieval, clearer communication, and stronger follow-through.
- Command high-severity incidents as Incident Commander and reinforce the systems, tools, and practices that simplify incident management.
- Design and implement workflows in which AI assists with alert correlation, signal enrichment, root-cause exploration, runbook retrieval, postmortem drafting, and corrective-action tracking.
On-call Quality and Toil Reduction
- Elevate on-call quality by silencing noisy alerts, automating repetitive investigations, and enabling responders to rapidly digest service context.
- Build tools that gather context from systems like Datadog, CloudWatch, incident.io, Slack, runbooks, deployment history, and service metadata.
- Transition teams from reactive paging to proactive reliability enhancement.
Architecture and Resilience
- Steer service designs for graceful degradation, failure isolation, robust capacity planning, and operational safety throughout EarnIn’s AWS environment.
- Apply production data, incident learnings, and AI analysis to spot architectural risks before they recur.
- Instruct engineering teams to embed reliability expectations into design reviews, launch protocols, and service evolution.
Mentorship and Cross-org Influence
- Coach engineers in reliability practices, incident response, SLOs, observability, production debugging, and AI-assisted operational workflows.
- Direct design reviews, incident reviews, and operational maturity discussions to improve engineering judgment across teams.
- Produce documentation, tooling, and reusable patterns that unlock reliability knowledge and enable action.
Requirements
- 7+ years in SRE, Software Engineering, or Infrastructure Engineering with increasing scope and cross-org influence.
- Track record of KPI-driven reliability and operational excellence improvements at scale.
- Demonstrated experience improving reliability and operational excellence at scale using clear KPIs.
- Shipped experience applying AI/LLMs to engineering or operational workflows.
- Significant expertise with SLIs, SLOs, error budgets, incident command, blameless postmortems, and recurrence prevention in large-scale distributed systems.
- Strong software engineering ability in Python, Go, or similar languages.
- Deep observability experience with systems such as Datadog, CloudWatch, OpenTelemetry, or similar platforms.
- Strong infrastructure-as-code and cloud infrastructure experience, including Terraform, Kubernetes, AWS, and safe, reversible deployment practices.
Benefits
- Base salary range: $252,000-$308,000, plus equity and benefits.
- Hybrid position in Mountain View (Headquarters) with in-office work 2 days a week.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/earnin/jobs/8193516