New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Okta

Senior Site Reliability Engineer

Okta
Apply →
hybrid senior full-time Bengaluru, India

First indexed 9 Sept 2026

Description

We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG).

The ideal candidate will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services.

Reliability & Operations

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in an on-call rotation supporting highly available customer-facing systems.
  • Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Partner with engineering teams to improve service availability, scalability, performance, and resilience.
  • Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.

Engineering & Automation

  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.
  • Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.

Technical Leadership

  • Contribute to and drive reliability initiatives within the product group.
  • Guide engineers in adopting operational best practices and reliability engineering principles.
  • Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
  • Support architecture and operational decisions through data-driven recommendations and engineering expertise.
  • Execute projects from conception through production rollout and long-term operational ownership.

Innovation

  • Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.
  • Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity.

Our Tech Stack

  • Infrastructure/Orchestration: Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, GitOps
  • Programming: Golang, Python
  • Observability: Datadog, Splunk
  • Data Stores: PostgreSQL, Redis, OpenSearch
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/okta/jobs/8078730