# Machine Learning Operations Developer

**Company**: Electronic Arts
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Salary**: $141,400 - $204,400 CAD
**Category**: Engineering
**Industry**: Technology
**Ticker**: EA
**Wikidata**: https://www.wikidata.org/wiki/Q173941

**Apply**: https://jobs.ea.com/fr_CA/careers/JobDetail/MLOps-Developer/216108?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_9fa9b10b-0d6

## Description

Electronic Arts is seeking a Machine Learning Operations Developer to work closely with researchers to optimise models, prepare required training data, and measure and improve model performance.

As a Machine Learning Operations Developer, you will work on optimising how our models are trained and served to users, creating a data pipeline from the data lake to graphics processing units, and ensuring intentional data format choices.

Responsibilities:

- Optimise model training and serving

- Create data pipelines from data lakes to graphics processing units

- Ensure intentional data format choices

- Instrument infrastructure to measure GPU usage, task queue lengths, task success rates, experimentation costs, data loading wait times, and inference latency percentiles

- Aggregate telemetry data related to training and experiments into a single repository

Requirements:

- At least 7 years of development experience

- At least 4 years of experience in designing and operating production machine learning systems

- Expertise in Python

- Ability to read and modify code at the software infrastructure level, including PyTorch

- Experience with inference servers like vLLM, SGLang, or TensorRT-LLM

- Proven experience in model optimisation or inference

- Practical experience with distributed data processing for machine learning (e.g., Ray, Spark)

- Familiarity with columnar formats for machine learning (e.g., Parquet, Arrow, Lance)

- Practical skills in observability (e.g., Prometheus, Grafana)

- Knowledge of AWS compute and storage services for machine learning workloads (e.g., EC2 instances with GPUs, S3, EKS)

- Emphasis on reproducibility and experience with containerisation and continuous integration/continuous deployment for machine learning artefacts

This is a hybrid role, requiring three days per week in one of our offices in Montreal, Redwood City, or Vancouver.

Salary ranges:

- British Columbia: $141,400 - $204,400 CAD

- California: $165,000 - $256,000 USD

Benefits include paid time off, sick time, company holidays, medical/dental/vision insurance, life insurance, disability insurance, and retirement plans.

## Skills

### Required
- Python
- PyTorch
- machine learning
- data processing
- distributed systems
- observability
- AWS
- containerisation
- continuous integration

---

Source: [Apply at jobs.ea.com](https://jobs.ea.com/fr_CA/careers/JobDetail/MLOps-Developer/216108?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
