# Live Support Engineer

**Company**: Electronic Arts
**Location**: Shanghai
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology
**Ticker**: EA
**Wikidata**: https://www.wikidata.org/wiki/Q173941

**Apply**: https://jobs.ea.com/es_ES/careers/JobDetail/216151-Live-Support-Engineer-II/216151?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_359c9cdf-e51

## Description

Electronic Arts creates incredible entertainment experiences that inspire players and fans worldwide. Here, everyone is part of the story, part of a community that connects people globally.

As one of the largest sports entertainment platforms in the world, EA SPORTS FC is redefining football with genre-leading interactive experiences, connecting a global community of fans to The World's Game through innovation and unrivalled authenticity.

You will report to the DevOps Director.

### Responsibilities

#### Prod and Pre-Prod Environment Monitoring

- Monitor client and server stability across pre-production and production environments, including crashes, ANRs, error rates, key performance metrics, service health, core gameplay flows, and infrastructure status.

- Use Firebase, Grafana, Sentry, Splunk, VictoriaLogs, and VictoriaMetrics to identify abnormal trends and correlate metrics, logs, and error events.

- Identify and drive improvements to alert noise, monitoring gaps, missing telemetry, and recurring issues, while continuously optimising dashboards and alerting rules.

- Provide stability support during release windows and major esports events.

#### Live Incident Response

- Work as a first-line responder for live incidents, following Incident Runbooks for investigation, communication, escalation, and recovery.

- Provide updates on impact, investigation progress, mitigation actions, and recovery status throughout an incident.

- Follow approval processes and Runbooks to perform mitigation actions such as restarts, failovers, configuration changes, and Kill Switch activation.

- Maintain complete operational records for post-incident review and audit purposes.

- Escalate promptly to service owners, development teams, or infrastructure teams in accordance with the Escalation Policy.

#### AI-Powered Live Monitoring and Operations Tools

- Build AI-powered live monitoring, alert analysis, and operations support tools, including internal AI Agents, MCP tools, knowledge bases, prompts, evaluation cases, and runtime configurations.

- Integrate data sources and APIs from monitoring, logging, and collaboration platforms to automatically collect and correlate metrics, logs, and error information.

- Turn repetitive health checks, anomaly summaries, impact analysis, and incident context collection into automated workflows.

- Evaluate the accuracy of AI-generated analysis and reduce false positives, false negatives, and non-applicable recommendations.

#### Knowledge and Process Development

- Develop knowledge bases covering client data flows, service architecture, key dependencies, error codes, and metric catalogues, along with dashboard guides and troubleshooting documentation.

- Convert incident insights into applicable Runbooks.

- Work with development teams to complete the handover and acceptance of dashboards, alerts, and Runbooks, and drive documentation updates after feature releases.

- Track incident action items and lead responsible teams to deliver long-term fixes.

### Qualifications

- A degree in Computer Science or a related field, or equivalent technical capability with three or more years of experience in SRE, operations, production support, or technical support roles.

- Familiarity with Kubernetes, containers, cloud platforms, and microservices architecture.

- Foundational development skills in Go, Python, or another programming language, with the ability to integrate APIs, write automation scripts, and maintain internal tools.

- Familiarity with Prometheus, ELK, Grafana, Sentry, Splunk, VictoriaLogs, VictoriaMetrics, and Firebase.

- Basic knowledge of large language models, AI Agents, prompting, RAG, or MCP, with the ability to validate the reliability of AI-generated output.

- Read technical documentation in English and communicate in writing with global development and operations teams.

- Willing to participate in shift rotations, release support, and on-call coverage.

### Benefits

Electronic Arts offers a comprehensive benefits package, including physical, emotional, financial, professional, and community well-being support. Benefits are customised to meet local needs and may include medical insurance, mental wellness support, pension plans, paid time off, family leave, free games, and more.

## Skills

### Required
- Kubernetes
- containers
- cloud platforms
- microservices architecture
- Go
- Python
- Prometheus
- ELK
- Grafana
- Sentry
- Splunk
- VictoriaLogs
- VictoriaMetrics
- Firebase

### Nice to have
- large language models
- AI Agents
- prompting
- RAG
- MCP

---

Source: [Apply at jobs.ea.com](https://jobs.ea.com/es_ES/careers/JobDetail/216151-Live-Support-Engineer-II/216151?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
