# HPC Network Engineer

**Company**: Fuse Energy
**Location**: London
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: IT
**Industry**: Energy

**Apply**: https://jobs.workable.com/view/dTK5ymsmbfzXo7z6j8VMRJ/hpc-network-engineer-in-london-at-fuse-energy?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_82fa4338-68e

## Description

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast.

The Opportunity You'll design, deploy, and operate the network fabric for our multi-tenant AI cluster. This covers the full stack: the high-performance compute and storage fabrics carrying RDMA traffic between GPUs, the tenant-facing and management networks, fire walling and tenant isolation, and the out-of-band infrastructure that keeps it all recoverable.

Responsibilities

- Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scale

- Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN)

- Implement and maintain per-tenant network isolation across compute, storage, and management planes

- Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI

- Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry)

- Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts

- Operate the out-of-band management network, console access, and remote recovery paths

- Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales

- Write clear design documentation capturing decisions, rationale, and rejected alternatives

- Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environments

- Upskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidently

Requirements

- 5+ years as a network engineer operating production data centre networks

- Strong dynamic routing experience (BGP in particular), plus overlay/encapsulation design and troubleshooting (e.g. EVPN/VXLAN)

- Hands-on experience with leaf-spine / Clos fabric design and operation

- Experience with modern data centre network operating systems and comfortable in the Linux networking stack, not just a vendor CLI

- Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, with an understanding of why lossless behaviour matters for GPU workloads

- Network automation as a working practice, not an aspiration: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source of truth, version-controlled changes

- Solid Linux administration fundamentals: you can debug from the host side as well as the switch side

- Experience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)

- Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access (e.g. Cisco Catalyst/Meraki or comparable)

- Clear communicator who enjoys teaching: able to document, pair, and run sessions that bring less network-savvy colleagues up to speed

Nice to have

- Experience with GPU cluster networking specifically (e.g. NVIDIA Spectrum-X or Quantum InfiniBand, ConnectX/BlueField NICs and DPUs, UFM, SHARP, or equivalent Broadcom/Arista AI fabric platforms)

- Container networking experience: CNI plugins and BGP integration between clusters and the fabric

- Understanding of collective-communication libraries and how fabric behaviour shows up as training/inference performance

- Multi-tenant network design: VRF-based isolation, tenant bandwidth guarantees, secure shared infrastructure

- Experience with enterprise firewall platforms (e.g. FortiGate, Palo Alto), including HA deployment and virtualised/segmented instances

- Storage networking experience (e.g. NVMe-oF, lossless storage fabrics, per-tenant storage isolation)

- Bare-metal provisioning environments (e.g. MAAS, PXE, Redfish) and how network bootstrap fits into node lifecycle

- Optical layer knowledge at 200/400/800G: transceivers, MPO cabling, link qualification

- Experience standing up a data centre network from greenfield

- Relevant certifications (e.g. CCNP/CCIE or equivalent), valued as evidence of depth, not a gate

Benefits

- Competitive salary and an equity sign-on bonus

- Biannual bonus scheme

- Fully expensed tech to match your needs

- Breakfast and dinner allowance for office-based employees

## Skills

### Required
- Network engineering
- Data centre networks
- Dynamic routing
- Overlay/encapsulation design
- Leaf-spine / Clos fabric design
- RDMA fabric experience
- Network automation
- Linux administration
- Network telemetry and monitoring
- Corporate/campus networks

### Nice to have
- GPU cluster networking
- Container networking
- Collective-communication libraries
- Multi-tenant network design
- Enterprise firewall platforms
- Storage networking
- Bare-metal provisioning environments
- Optical layer knowledge

---

Source: [Apply at jobs.workable.com](https://jobs.workable.com/view/dTK5ymsmbfzXo7z6j8VMRJ/hpc-network-engineer-in-london-at-fuse-energy?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
