Description
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.
We are seeking a Network Engineer who is strong on fundamentals and comfortable owning production network design, deployment, and operations end to end. This is a hands-on engineering seat , not a NOC technician role and not a network-software (telemetry/ZTP platform) SWE role.
Responsibilities:
- Design, deploy, and operate production datacenter and campus/core networks at scale
- Own routing and switching configuration standards (BGP and at least one IGP such as OSPF or IS-IS), including change design, peer reviews, and execution
- Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
- Build and improve monitoring, alerting, and operational documentation so issues are caught and fixed quickly
- Troubleshoot Layer 2/Layer 3 incidents end to end , from link flaps and optics through routing and traffic engineering , and drive root cause and lasting fixes
- Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil
- Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows
- Support high-performance / supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA-capable designs) as part of the broader network estate , deep specialist RoCE/NCCL ownership is a plus, not the bar for this seat
Basic Qualifications:
- Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
- Solid hands-on experience with BGP and at least one interior routing protocol
- Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high-speed Ethernet
- Experience troubleshooting live production network incidents and participating in on-call
- Strong written and verbal communication; clear change docs and incident notes
Preferred Skills and Experience:
- Experience with modern datacenter vendors (e.g. Arista, Cisco, Juniper, Nvidia/Mellanox)
- Familiarity with high-performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics) , useful context for our environment, not a hard filter
- Network automation (Python, Ansible, Terraform, or similar) used in production
- Experience with EVPN, leaf-spine, and large-scale Ethernet fabrics
- Prior work supporting rapid datacenter or cluster capacity build-outs
Additional Requirements:
- Willing to work onsite in Palo Alto
Compensation and Benefits:
$150,000 - $250,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/xai/jobs/5235761007