Description
We are looking for a technical leader to oversee the Site Reliability Engineering (SRE) organisation, focusing on Okta's platform, databases, edge networking, Kubernetes, CI/CD, observability, FinOps, and automation.
As the Director of Site Reliability Engineering, you will:
- Build and lead a high-calibre India-based SRE organisation supporting Okta's production fleet.
- Partner with global engineering, product, and infrastructure leaders to deliver resilient, scalable, and secure services.
- Define and execute the India SRE strategy in alignment with global reliability goals.
- Lead post-incident reviews, drive root-cause analysis, and ensure long-term corrective actions.
- Implement automation and observability to reduce manual toil and improve operational efficiency.
- Hire, mentor, and develop top SRE talent across India.
Requirements:
- 16+ years of experience in site reliability, infrastructure, or production engineering roles.
- 8+ years of experience in technical leadership and people management.
- Strong expertise in automation, observability, performance optimisation, and incident response.
- Experience building or scaling offshore SRE teams.
- 4+ years of experience running the SRE org supporting a SaaS/Cloud service in a public Cloud, preferably AWS.
- Strong expertise in cloud-native architectures, containerisation, IaC, and CI/CD pipelines.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/okta/jobs/8155879