Description
Electronic Arts is seeking an experienced Senior Site Reliability Engineer (SRE) to lead the reliability, scalability, and performance engineering of critical infrastructure and production systems.
As a Senior SRE, you will play a strategic and technical leadership role, driving reliability practices, mentoring SRE teams, and influencing the adoption of automation, observability, and resilience engineering across the organisation.
Responsibilities
Reliability Engineering & Automation
- Design, implement, and manage resilient, scalable, and highly available infrastructure systems.
- Lead initiatives to automate manual operations, deployment, and monitoring processes to improve reliability and reduce toil.
- Drive the creation of observability solutions and dashboards to proactively detect and remediate potential issues.
Incident & Problem Management
- Lead critical incident response, ensuring swift mitigation and clear communication to stakeholders.
- Conduct detailed root cause analysis (RCA) and drive permanent corrective actions to prevent recurrence.
- Implement and mature incident management frameworks, including runbooks, playbooks, and post-incident reviews.
Infrastructure Operations & Performance Optimisation
- Oversee system performance, capacity planning, and scalability of infrastructure across hybrid and cloud environments (AWS, Azure, GCP).
- Optimise system resource utilisation, latency, and reliability through performance tuning and automation.
- Work closely with architecture and platform teams to accommodate growth, change, and modernisation initiatives.
Leadership & Mentorship
- Provide technical leadership and mentorship to SRE teams and cross-functional engineering groups.
- Promote an SRE culture across teams, championing principles of reliability, automation, observability, and continuous improvement.
- Drive collaboration between development, QA, DevOps, and release teams to embed reliability into the software development lifecycle (SDLC).
Service Level Management
- Define, track, and continuously improve Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- Apply the Four Golden Signals of SRE monitoring , Latency, Traffic, Errors, and Saturation , to guide system health and performance strategies.
Documentation & Knowledge Sharing
- Establish and maintain comprehensive documentation of systems, operational procedures, and best practices.
- Facilitate learning through technical sessions, blameless postmortems, and cross-team knowledge sharing.
Strategic Technology & Continuous Improvement
- Contribute to defining the long-term SRE strategy, tooling roadmap, and automation frameworks.
- Evaluate and adopt emerging technologies, tools, and methodologies to enhance system reliability and efficiency.
- Partner with business and technical leaders to ensure alignment of SRE objectives with organisational goals.
Security & Compliance
- Collaborate with security and compliance teams to ensure infrastructure, systems, and operations meet organisational and regulatory standards.
- Implement secure configuration baselines, vulnerability remediation, and access control policies.
- Integrate security practices into CI/CD pipelines to ensure DevSecOps alignment.
Strategic Leadership & Stakeholder Management
- Partner with executive and business stakeholders to align SRE initiatives with enterprise objectives and risk frameworks.
- Provide data-driven insights on reliability, capacity, and operational performance to influence strategic decision-making.
- Represent SRE functions in technical governance forums, audits, and architecture reviews to drive reliability-focused outcomes.
Requirements
- Education: Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.
- Experience: 12–15 years of total IT experience, with at least 8+ years in SRE, DevOps, or large-scale systems engineering.
- Technical Expertise:
- Strong proficiency in Linux/Unix system administration and internals.
- Proven experience in cloud platforms , AWS, Azure, or GCP.
- Advanced scripting and automation skills using Python, Go, PowerShell, or Bash.
- Hands-on exposure to containerisation and orchestration technologies (Docker, Kubernetes) and expertise on service mesh like Istio etc.
- Deep understanding of monitoring and observability stacks (Prometheus, Grafana, ELK, Datadog, Splunk, Zabbix, Nagios).
- Expertise in configuration management and IaC tools (Ansible, Terraform, Chef, Puppet).
- Strong knowledge of networking, load balancing, databases, and distributed systems.
- Operational Excellence:
- Hands-on experience in incident response, problem management, and capacity planning at enterprise scale.
- Proven ability to design for reliability, redundancy, and disaster recovery.
- Soft Skills:
- Excellent analytical, communication, and leadership abilities.
- Proven track record of mentoring and developing high-performing engineering teams.
- Strong stakeholder management and cross-functional collaboration skills.
Nice to Have
- Experience defining and implementing SRE frameworks or centres of excellence in global organisations.
- Familiarity with REST API development, integration, and database query optimisation.
- Strong understanding of governance, risk, and compliance frameworks.
- Experience with AIOps, self-healing systems, or machine learning-driven monitoring.
- Demonstrated experience in driving organisational culture change toward reliability and automation.
- Active participation in industry forums or open-source contributions related to DevOps or SRE practices.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://jobs.ea.com/es_ES/careers/JobDetail/Senior-Site-Reliability-Engineer-I/211515