Description
SpaceXAI is seeking a Network Engineer to design, deploy, and operate large-scale networks that support training and inference infrastructure. The ideal candidate will have strong fundamentals, own production network design, deployment, and operations end-to-end, and be comfortable with hands-on engineering.
Responsibilities:
- Design, deploy, and operate production datacenter and campus/core networks at scale
- Own routing and switching configuration standards, including change design, peer reviews, and execution
- Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
- Build and improve monitoring, alerting, and operational documentation
- Troubleshoot Layer 2/Layer 3 incidents end-to-end
- Automate repetitive network tasks with Python, Ansible, or similar tooling
- Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows
- Support high-performance/supercompute network environments
Basic Qualifications:
- Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
- Solid hands-on experience with BGP and at least one interior routing protocol
- Working knowledge of TCP/IP, VLANs, EVPN/VXLAN, and optics/high-speed Ethernet
- Experience troubleshooting live production network incidents and participating in on-call
- Strong written and verbal communication
Preferred Skills and Experience:
- Experience with modern datacenter vendors (e.g., Arista, Cisco, Juniper, Nvidia/Mellanox)
- Familiarity with high-performance or supercompute networking
- Network automation (Python, Ansible, Terraform, or similar) used in production
- Experience with EVPN, leaf-spine, and large-scale Ethernet fabrics
- Prior work supporting rapid datacenter or cluster capacity build-outs
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/xai/jobs/5235761007