Description
As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform , the secure, high-performance code execution layer powering our agentic workflows, deployed across both internal and customer-managed environments.
This is a role for someone who cares as much about the experience of the engineers and researchers using this system as they do about the kernel internals underneath it.
You will:
- Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments
- Ensure strong isolation, security, and reproducibility of execution across user sessions and workloads
- Optimise for cold-start latency, memory footprint, and resource utilisation at scale
- Drive down error rates through systematic debugging, monitoring, and proactive fixes
- Partner closely with internal teams using the platform to understand their needs, debug issues, and build tooling that serves their use cases
- Respond to incidents and production issues with urgency, conducting root cause analysis and implementing preventive fixes
- Help develop and maintain a product roadmap for sandboxing, balancing immediate needs against long-term architectural investment
- Lead architecture reviews and own projects end-to-end, from design through deployment, in fast-paced cross-functional settings
Ideally you'd have:
- 4+ years of experience building high-performance systems software, with meaningful time spent maintaining libraries, SDKs, or developer-facing APIs
- Deep understanding of Linux internals: process isolation, memory management, cgroups, namespaces, etc.
- Experience with containerisation and virtualisation technologies (e.g., Docker, Firecracker, gVisor, QEMU, Kata Containers)
- Proficiency in a systems programming language such as Go, Rust, or C/C++
- A track record of obsessing over developer experience , API design, error propagation, documentation, and the small details that make a library feel well-crafted
- Comfort working across infrastructure layers, from kernel modules to orchestration frameworks (e.g., Kubernetes)
- Strong debugging skills and the ability to navigate performance/security tradeoffs in production systems
- Comfort with ambiguity, and the ability to context-switch between reactive incident work and proactive product development
Nice to haves:
- Experience as a founder or early engineer at an infrastructure-focused startup, owning a product end-to-end
- Familiarity with LLM agents and agent frameworks (e.g., OpenHands, Agent2Agent, MCP)
- Experience running secure workloads in multi-tenant or untrusted environments (e.g., FaaS, CI sandboxes, remote notebooks)
- Exposure to snapshotting and restore techniques (e.g., CRIU, VM snapshots, overlays)
- Open-source contributions to systems or developer-tools projects
- History of on-call/incident response for production systems
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/scaleai/jobs/4717105005