# Senior Site Reliability Engineer

**Company**: Okta
**Location**: Bengaluru, India
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/okta/jobs/8078730?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_3da2a2d5-7b8

## Description

We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG).

The ideal candidate will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services.

**Reliability & Operations**

- Design, build, and operate large-scale cloud infrastructure and production services.

- Participate in an on-call rotation supporting highly available customer-facing systems.

- Lead incident response efforts and drive post-incident reviews focused on systemic improvements.

- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

- Partner with engineering teams to improve service availability, scalability, performance, and resilience.

- Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.

**Engineering & Automation**

- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.

- Eliminate operational toil through automation, tooling, and platform engineering.

- Improve deployment safety and operational workflows through CI/CD and GitOps practices.

- Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.

- Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.

**Technical Leadership**

- Contribute to and drive reliability initiatives within the product group.

- Guide engineers in adopting operational best practices and reliability engineering principles.

- Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.

- Support architecture and operational decisions through data-driven recommendations and engineering expertise.

- Execute projects from conception through production rollout and long-term operational ownership.

**Innovation**

- Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.

- Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity.

**Our Tech Stack**

- Infrastructure/Orchestration: Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, GitOps

- Programming: Golang, Python

- Observability: Datadog, Splunk

- Data Stores: PostgreSQL, Redis, OpenSearch

## Skills

### Required
- Kubernetes
- Terraform
- Go
- Python
- Infrastructure as Code
- Cloud Networking
- Observability Platforms
- Distributed Data Platforms

### Nice to have
- Experience operating SaaS platforms serving large-scale customer workloads
- Experience working within Kubernetes-based microservices environments
- Experience supporting globally distributed production environments
- Experience with GitOps and ArgoCD
- Experience implementing AI-assisted operational tooling or automation workflows

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/okta/jobs/8078730?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
