# Senior Site Reliability Engineer

**Company**: Kody
**Location**: Hong Kong
**Experience**: senior
**Category**: Engineering
**Industry**: Finance

**Apply**: https://jobs.workable.com/view/d5EHEnd84nFiZytV8DsQ3p/senior-site-reliability-engineer-in-hong-kong-at-kody?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_ede8a9ac-bed

## Description

Kody is seeking a Senior Site Reliability Engineer to drive the reliability, availability, scalability, and operational excellence of its global payment platform.

Key Responsibilities:

- Participate in a follow-the-sun production on-call rotation as a senior incident responder.

- Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.

- Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.

- Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.

- Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).

- Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.

- Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

Requirements:

- 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.

- Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).

- Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.

- Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.

- Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.

- Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

Benefits:

- Competitive Package

- A dynamic and innovative team

- Collaborative, inclusive working environment

## Skills

### Required
- AWS
- Kubernetes
- Terraform
- PostgreSQL
- Redis
- Kafka
- Linux
- Datadog
- Prometheus
- Grafana
- PCI-DSS

---

Source: [Apply at jobs.workable.com](https://jobs.workable.com/view/d5EHEnd84nFiZytV8DsQ3p/senior-site-reliability-engineer-in-hong-kong-at-kody?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
