# Senior Machine Learning Engineer

**Company**: Cloudflare
**Location**: Austin, TX
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/cloudflare/jobs/8043974?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_5f67930c-642

## Description

At Cloudflare, we're on a mission to help build a better Internet.

You'll help define how machine learning models run across Cloudflare's global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators.

## Responsibilities

- Develop, optimize, and productionize machine learning models for Cloudflare's serverless inference platform, with a focus on performance, reliability, and model quality.

- Build benchmarking and evaluation frameworks to measure latency, throughput, cost efficiency, and model behavior across LLMs, speech, vision, and other model families.

- Improve inference performance through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization.

- Partner with systems engineers to integrate models into Cloudflare's distributed inference infrastructure across a heterogeneous fleet of GPUs and next-generation accelerators.

- Drive improvements to model deployment workflows, including validation, rollout safety, observability, regression testing, and operational readiness.

- Collaborate with product and engineering teams to translate customer requirements into scalable ML capabilities for Workers AI.

- Mentor engineers, contribute to technical direction, and raise the quality bar for production ML engineering practices across the team.

## Desirable Skills, Knowledge, and Experience

- Experience building, optimizing, and operating machine learning models in production environments.

- Strong proficiency with Python and modern ML frameworks such as PyTorch, TensorFlow, JAX, or equivalent.

- Hands-on experience with inference optimization techniques for large-scale models, including quantization, batching, caching, compilation, and serving runtime tuning.

- Experience with large-scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar.

- Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures.

- Experience optimizing models for GPUs or specialized accelerators.

- Strong understanding of production ML concerns, including evaluation, monitoring, model regressions, rollout safety, and reliability.

- Ability to work across ML and systems boundaries, including familiarity with distributed systems, networking, or serverless platforms.

- Track record of leading complex technical projects and mentoring other engineers.

## Skills

### Required
- Python
- PyTorch
- TensorFlow
- JAX
- machine learning
- inference optimization
- large-scale models
- distributed systems
- serverless platforms

### Nice to have
- experience with open source ML tooling
- model serving frameworks
- inference runtimes

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/cloudflare/jobs/8043974?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
