# AI Inference Engineer

**Company**: Fuse Energy
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://jobs.workable.com/view/d8L9Rw7VMr8fCNDwZqPaQ6/remote-ai-inference-engineer-in-united-states-at-fuse-energy?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_0e33011f-86a

## Description

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast.

We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.

## Responsibilities

- Define Fuse's inference serving strategy and architecture from first principles.

- Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.

- Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.

- Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).

- Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.

- Act as a direct technical owner of inference performance and reliability.

- Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.

- Set the standards, tooling, and benchmarks this function will run on as it grows.

## Requirements

- 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.

- Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).

- Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.

- Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.

- A track record of making high-stakes architecture calls and owning the outcome.

- Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.

## Nice to Have

- Experience with Triton or custom ML inference/training frameworks.

- Experience with autoscaling or capacity planning for large-scale inference workloads.

- Exposure to multi-tenant serving or SLA-driven infrastructure.

- Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.

- Familiarity with Kubernetes/Slurm for cluster orchestration.

- Interest or experience in energy markets, grid systems, or sustainability-focused compute.

## Benefits

- Competitive salary and an equity sign-on bonus.

- Biannual bonus scheme.

- Fully expensed tech to match your needs.

- Breakfast and dinner allowance for office-based employees.

## Skills

### Required
- large-scale inference serving systems
- inference serving frameworks
- systems thinking
- GPU/CUDA engineering
- architecture calls

### Nice to have
- Triton
- custom ML inference/training frameworks
- autoscaling
- capacity planning
- multi-tenant serving
- SLA-driven infrastructure
- Kubernetes/Slurm

---

Source: [Apply at jobs.workable.com](https://jobs.workable.com/view/d8L9Rw7VMr8fCNDwZqPaQ6/remote-ai-inference-engineer-in-united-states-at-fuse-energy?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
