Description
Electronic Arts creates next-level entertainment experiences that inspire players and fans worldwide. As a Site Reliability Engineer III, you will cover the entire lifecycle of a product, from helping developers with architecture and delivery to on-call incident response and triage.
Responsibilities:
- Build and operate distributed, large-scale, cloud-based infrastructure using modern open-source software solutions.
- Use automation technologies to ensure repeatability, eliminate toil, reduce mean time to detection and resolution (MTTD & MTTR), and repair services.
- Perform root cause analysis and post-mortems with an eye towards future prevention.
- Design and build CI/CD pipelines.
- Create monitoring, alerting, and dashboarding solutions that improve visibility into EA's application performance and business metrics.
- Produce documentation and support tooling for online support teams.
Qualifications:
- 7+ years of experience with Virtualization, Containerization, Cloud Computing (AWS preferred), Kubernetes, or Docker.
- 7+ years of experience supporting high-availability production-grade infrastructure and applications with defined SLIs and SLOs.
- Systems Administration experience, including a strong understanding of Linux.
- Network experience, including an understanding of standard protocols/components.
- Automation and orchestration experience including Terraform, Helm, Chef, Puppet, Packer.
- Experience writing code in Python, Golang, or Java.
- Experience working with distributed systems.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://jobs.ea.com/ms_MY/careers/JobDetail/Site-Reliability-Engineer-III/215708