# Site Reliability Engineer II

**Company**: Electronic Arts
**Location**: Hyderabad, Telangana
**Work arrangement**: hybrid
**Experience**: mid
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology
**Ticker**: EA
**Wikidata**: https://www.wikidata.org/wiki/Q173941

**Apply**: https://jobs.ea.com/fr_CA/careers/JobDetail/Site-Reliablity-Engineer-II/216206?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_cacc3301-a49

## Description

As a Site Reliability Engineer II at Electronic Arts, you will contribute to designing, automating, and operating large-scale, cloud-based systems that power EA's global gaming platform.

**Responsibilities:**

- Build and operate scalable systems: Support development, deployment, and maintenance of distributed, cloud-based infrastructure using modern open-source technologies (AWS/GCP/Azure, Kubernetes, Terraform, Docker, etc.).

- Platform operations and automation: Develop automation scripts, tools, and workflows to reduce manual effort, improve system reliability, and optimise infrastructure operations.

- Monitoring, alerting, and incident response: Create and maintain dashboards, alerts, and metrics to improve system visibility and proactively identify issues.

- Continuous Integration/Continuous Deployment (CI/CD): Contribute to designing, implementing, and maintaining CI/CD pipelines for consistent, repeatable, and reliable deployments.

- Reliability and performance engineering: Collaborate with cross-functional teams to identify reliability bottlenecks, define SLIs/SLOs/SLAs, and implement improvements.

- Post-incident reviews and documentation: Participate in root cause analyses, document learnings, and contribute to preventive measures.

- Collaboration and mentorship: Work closely with senior SREs and software engineers to gain exposure to large-scale systems and adopt best practices.

- Modernisation and continuous improvement: Contribute to ongoing modernisation efforts by identifying areas for improvement in automation, monitoring, and reliability.

**Qualifications:**

- 3–5 years of experience in Cloud Computing (AWS preferred), Virtualization, and Containerization using Kubernetes, Docker, or VMWare.

- Extensive hands-on experience in container orchestration technologies, such as EKS, Kubernetes, Docker.

- Experience supporting production-grade, high-availability systems with defined SLIs/SLOs.

- Strong Linux/Unix administration and networking fundamentals (protocols, load balancing, DNS, firewalls).

- Hands-on experience with Infrastructure as Code and automation tools such as Terraform, Helm, Ansible, or Chef.

- Proficiency in Python, Golang, Bash, or Java for scripting and automation.

- Familiar with monitoring and observability tools like Prometheus, Grafana, Loki, or Datadog.

- Exposure to distributed systems, SQL/NoSQL databases, and CI/CD pipelines.

- Strong problem-solving, troubleshooting, and collaboration skills in cross-functional environments.

**Benefits:**

- Comprehensive benefits programs focusing on physical, emotional, financial, professional, and community well-being.

- Programs designed to meet local needs, including medical coverage, mental well-being support, retirement savings, paid leaves, parental leaves, free games, and more.

## Skills

### Required
- Cloud Computing
- Virtualization
- Containerization
- Kubernetes
- Docker
- Terraform
- Python
- Golang
- Bash
- Java
- Prometheus
- Grafana
- Loki
- Datadog

### Nice to have
- AWS
- GCP
- Azure
- EKS
- Ansible
- Chef
- Helm

---

Source: [Apply at jobs.ea.com](https://jobs.ea.com/fr_CA/careers/JobDetail/Site-Reliablity-Engineer-II/216206?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
