Description
Electronic Arts is seeking a Site Reliability Engineer III to join their Production Infrastructure & Engineering (PI&E) organization. As a Site Reliability Engineer, you will cover the entire lifecycle of a product, from helping developers with architecture and delivery to on-call incident response and triage.
Your primary focus will be automation and continuous integration/delivery with an emphasis on solving operations issues using software. You will report to the Senior SRE Manager.
Responsibilities:
- Build and operate distributed, large-scale, cloud-based infrastructure using modern open-source software solutions.
- Use automation technologies to ensure repeatability, eliminate toil, reduce mean time to detection and resolution (MTTD & MTTR) and repair services.
- Perform root cause analysis and post-mortems with an eye towards future prevention.
- Design and build CI/CD pipelines.
- Create monitoring, alerting and dashboarding solutions that improve visibility into EA's application performance and business metrics.
- Produce documentation and support tooling for online support teams.
Qualifications:
- 7+ years of experience with Virtualization, Containerization, Cloud Computing (AWS preferred), Kubernetes, or Docker.
- 7+ years of experience supporting high-availability production-grade infrastructure and applications with defined SLIs and SLOs.
- Systems Administration experience, including a strong understanding of Linux.
- Network experience, including an understanding of standard protocols/components.
- Automation and orchestration experience including Terraform, Helm, Chef, Puppet, Packer.
- Experience writing code in Python, Golang, or Java.
- Experience working with distributed systems.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://jobs.ea.com/ko_KR/careers/JobDetail/Site-Reliability-Engineer-III/215708