# Senior Site Reliability Engineer I

**Company**: Electronic Arts
**Location**: Hyderabad, Telangana
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: IT
**Industry**: Technology
**Ticker**: EA
**Wikidata**: https://www.wikidata.org/wiki/Q173941

**Apply**: https://jobs.ea.com/es_ES/careers/JobDetail/Senior-Site-Reliability-Engineer-I/211515?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_7581ab14-4f1

## Description

Electronic Arts is seeking an experienced Senior Site Reliability Engineer (SRE) to lead the reliability, scalability, and performance engineering of critical infrastructure and production systems.

As a Senior SRE, you will play a strategic and technical leadership role, driving reliability practices, mentoring SRE teams, and influencing the adoption of automation, observability, and resilience engineering across the organisation.

**Responsibilities**

### Reliability Engineering & Automation

- Design, implement, and manage resilient, scalable, and highly available infrastructure systems.

- Lead initiatives to automate manual operations, deployment, and monitoring processes to improve reliability and reduce toil.

- Drive the creation of observability solutions and dashboards to proactively detect and remediate potential issues.

### Incident & Problem Management

- Lead critical incident response, ensuring swift mitigation and clear communication to stakeholders.

- Conduct detailed root cause analysis (RCA) and drive permanent corrective actions to prevent recurrence.

- Implement and mature incident management frameworks, including runbooks, playbooks, and post-incident reviews.

### Infrastructure Operations & Performance Optimisation

- Oversee system performance, capacity planning, and scalability of infrastructure across hybrid and cloud environments (AWS, Azure, GCP).

- Optimise system resource utilisation, latency, and reliability through performance tuning and automation.

- Work closely with architecture and platform teams to accommodate growth, change, and modernisation initiatives.

### Leadership & Mentorship

- Provide technical leadership and mentorship to SRE teams and cross-functional engineering groups.

- Promote an SRE culture across teams, championing principles of reliability, automation, observability, and continuous improvement.

- Drive collaboration between development, QA, DevOps, and release teams to embed reliability into the software development lifecycle (SDLC).

### Service Level Management

- Define, track, and continuously improve Service Level Objectives (SLOs) and Service Level Indicators (SLIs).

- Apply the Four Golden Signals of SRE monitoring , Latency, Traffic, Errors, and Saturation , to guide system health and performance strategies.

### Documentation & Knowledge Sharing

- Establish and maintain comprehensive documentation of systems, operational procedures, and best practices.

- Facilitate learning through technical sessions, blameless postmortems, and cross-team knowledge sharing.

### Strategic Technology & Continuous Improvement

- Contribute to defining the long-term SRE strategy, tooling roadmap, and automation frameworks.

- Evaluate and adopt emerging technologies, tools, and methodologies to enhance system reliability and efficiency.

- Partner with business and technical leaders to ensure alignment of SRE objectives with organisational goals.

### Security & Compliance

- Collaborate with security and compliance teams to ensure infrastructure, systems, and operations meet organisational and regulatory standards.

- Implement secure configuration baselines, vulnerability remediation, and access control policies.

- Integrate security practices into CI/CD pipelines to ensure DevSecOps alignment.

### Strategic Leadership & Stakeholder Management

- Partner with executive and business stakeholders to align SRE initiatives with enterprise objectives and risk frameworks.

- Provide data-driven insights on reliability, capacity, and operational performance to influence strategic decision-making.

- Represent SRE functions in technical governance forums, audits, and architecture reviews to drive reliability-focused outcomes.

**Requirements**

- Education: Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.

- Experience: 12–15 years of total IT experience, with at least 8+ years in SRE, DevOps, or large-scale systems engineering.

- Technical Expertise:

- Strong proficiency in Linux/Unix system administration and internals.

- Proven experience in cloud platforms , AWS, Azure, or GCP.

- Advanced scripting and automation skills using Python, Go, PowerShell, or Bash.

- Hands-on exposure to containerisation and orchestration technologies (Docker, Kubernetes) and expertise on service mesh like Istio etc.

- Deep understanding of monitoring and observability stacks (Prometheus, Grafana, ELK, Datadog, Splunk, Zabbix, Nagios).

- Expertise in configuration management and IaC tools (Ansible, Terraform, Chef, Puppet).

- Strong knowledge of networking, load balancing, databases, and distributed systems.

- Operational Excellence:

- Hands-on experience in incident response, problem management, and capacity planning at enterprise scale.

- Proven ability to design for reliability, redundancy, and disaster recovery.

- Soft Skills:

- Excellent analytical, communication, and leadership abilities.

- Proven track record of mentoring and developing high-performing engineering teams.

- Strong stakeholder management and cross-functional collaboration skills.

**Nice to Have**

- Experience defining and implementing SRE frameworks or centres of excellence in global organisations.

- Familiarity with REST API development, integration, and database query optimisation.

- Strong understanding of governance, risk, and compliance frameworks.

- Experience with AIOps, self-healing systems, or machine learning-driven monitoring.

- Demonstrated experience in driving organisational culture change toward reliability and automation.

- Active participation in industry forums or open-source contributions related to DevOps or SRE practices.

## Skills

### Required
- Linux/Unix system administration
- Cloud platforms (AWS, Azure, GCP)
- Scripting and automation (Python, Go, PowerShell, Bash)
- Containerisation and orchestration (Docker, Kubernetes)
- Monitoring and observability stacks (Prometheus, Grafana, ELK, Datadog, Splunk, Zabbix, Nagios)
- Configuration management and IaC tools (Ansible, Terraform, Chef, Puppet)
- Networking
- Load balancing
- Databases
- Distributed systems

### Nice to have
- SRE frameworks or centres of excellence
- REST API development
- Database query optimisation
- Governance, risk, and compliance frameworks
- AIOps
- Self-healing systems
- Machine learning-driven monitoring

---

Source: [Apply at jobs.ea.com](https://jobs.ea.com/es_ES/careers/JobDetail/Senior-Site-Reliability-Engineer-I/211515?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
