New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
OpenAI

Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI
Apply →
hybrid senior Full time London

First indexed 10 Jul 2026

Description

About the Team

ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance.

As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.

This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI.

About the Role

We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure.

You will design and build the systems that manage GPU clusters at scale,from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization.

This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI.

Responsibilities

  • Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.
  • Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.
  • Improve observability, reliability, and operational efficiency across thousands of GPUs.
  • Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
  • Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.
  • Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform.
  • Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.

Requirements

  • 5+ years of software engineering experience building production infrastructure.
  • Strong programming skills in Go, Python, C++, Rust, or similar systems languages.
  • Experience designing and operating highly available distributed systems.
  • Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
  • Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
  • Excellent debugging, systems design, and operational problem-solving skills.
  • Strong communication skills and experience collaborating across engineering organizations.

Benefits

OpenAI is an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://jobs.ashbyhq.com/openai/f8b84ae5-743b-41c9-8432-02dff9993d6b