New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Scale

AI Infrastructure Engineer, Sandbox Platform

Scale
Apply →
senior full-time London, UK

First indexed 24 Jul 2026

Description

As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform , the secure, high-performance code execution layer powering our agentic workflows, deployed across both internal and customer-managed environments.

This is a role for someone who cares as much about the experience of the engineers and researchers using this system as they do about the kernel internals underneath it.

You will:

  • Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments
  • Ensure strong isolation, security, and reproducibility of execution across user sessions and workloads
  • Optimise for cold-start latency, memory footprint, and resource utilisation at scale
  • Drive down error rates through systematic debugging, monitoring, and proactive fixes
  • Partner closely with internal teams using the platform to understand their needs, debug issues, and build tooling that serves their use cases
  • Respond to incidents and production issues with urgency, conducting root cause analysis and implementing preventive fixes
  • Help develop and maintain a product roadmap for sandboxing, balancing immediate needs against long-term architectural investment
  • Lead architecture reviews and own projects end-to-end, from design through deployment, in fast-paced cross-functional settings

Ideally you'd have:

  • 4+ years of experience building high-performance systems software, with meaningful time spent maintaining libraries, SDKs, or developer-facing APIs
  • Deep understanding of Linux internals: process isolation, memory management, cgroups, namespaces, etc.
  • Experience with containerisation and virtualisation technologies (e.g., Docker, Firecracker, gVisor, QEMU, Kata Containers)
  • Proficiency in a systems programming language such as Go, Rust, or C/C++
  • A track record of obsessing over developer experience , API design, error propagation, documentation, and the small details that make a library feel well-crafted
  • Comfort working across infrastructure layers, from kernel modules to orchestration frameworks (e.g., Kubernetes)
  • Strong debugging skills and the ability to navigate performance/security tradeoffs in production systems
  • Comfort with ambiguity, and the ability to context-switch between reactive incident work and proactive product development

Nice to haves:

  • Experience as a founder or early engineer at an infrastructure-focused startup, owning a product end-to-end
  • Familiarity with LLM agents and agent frameworks (e.g., OpenHands, Agent2Agent, MCP)
  • Experience running secure workloads in multi-tenant or untrusted environments (e.g., FaaS, CI sandboxes, remote notebooks)
  • Exposure to snapshotting and restore techniques (e.g., CRIU, VM snapshots, overlays)
  • Open-source contributions to systems or developer-tools projects
  • History of on-call/incident response for production systems
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/scaleai/jobs/4717105005