# Staff Site Reliability Engineer

**Company**: Okta
**Location**: Bengaluru, India
**Work arrangement**: hybrid
**Experience**: staff
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/okta/jobs/8059472?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_26d9d9b3-b3f

## Description

## Job Summary

We are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). As a Staff Site Reliability Engineer, you will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services.

## Responsibilities

### Reliability & Operations

- Design, build, and operate large-scale cloud infrastructure and production services.

- Participate in an on-call rotation supporting highly available customer-facing systems.

- Lead incident response efforts and drive post-incident reviews focused on systemic improvements.

- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

- Partner with engineering teams to improve service availability, scalability, performance, and resilience.

- Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.

### Engineering & Automation

- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.

- Eliminate operational toil through automation, tooling, and platform engineering.

- Improve deployment safety and operational workflows through CI/CD and GitOps practices.

- Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.

- Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.

### Technical Leadership

- Lead complex reliability initiatives spanning multiple engineering teams.

- Guide engineers in adopting operational best practices and reliability engineering principles.

- Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.

- Influence architecture and operational decisions through data-driven recommendations and engineering expertise.

- Drive projects from conception through production rollout and long-term operational ownership.

### Innovation

- Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.

- Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity.

## Requirements

### Technical Excellence

- Strong experience operating large-scale production services in AWS and/or GCP.

- Deep expertise with Kubernetes in production environments.

- Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.

- Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.

- Strong software engineering skills in Golang and/or Python.

- Experience building automation and internal engineering platforms.

- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.

- Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management.

- Experience with observability platforms, monitoring strategies, and production telemetry.

- Experience with or strong interest in AI-assisted engineering and operational automation.

### Operational Excellence

- Strong expertise operating customer-facing production systems.

- Experience leading incident response and driving operational improvements.

- Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning.

- Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices.

- Proven ability to balance reliability, scalability, security, and engineering velocity.

### Security & Compliance

- Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design.

- Experience implementing operational controls and best practices in regulated or security-sensitive environments is a plus.

### Leadership

- Demonstrated success leading complex engineering initiatives across multiple teams.

- Strong collaboration and communication skills.

- Experience working effectively within globally distributed engineering organizations spanning multiple timezones and cultures.

- Experience mentoring engineers and elevating technical capabilities within an organization.

- Ability to influence technical direction through expertise, partnership, and execution.

## Preferred Qualifications

- Experience operating SaaS platforms serving large-scale customer workloads.

- Experience working within Kubernetes-based microservices environments.

- Experience supporting globally distributed production environments.

- Experience with GitOps and ArgoCD.

- Experience implementing AI-assisted operational tooling or automation workflows.

## Skills

### Required
- Kubernetes
- Terraform
- Helm
- Go
- Python
- PostgreSQL
- Redis
- OpenSearch
- AWS
- GCP
- Cloud Security
- IAM
- CI/CD
- GitOps

### Nice to have
- SaaS platforms
- Kubernetes-based microservices
- Globally distributed production environments
- ArgoCD
- AI-assisted operational tooling

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/okta/jobs/8059472?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
