Description
SpaceXAI is seeking a Network Engineer to support the design, build-out, and operation of networks powering AI supercomputer campuses. The ideal candidate has experience in mission-critical, large-scale production environments.
Responsibilities:
- Design and implement highly available, low-latency, high-bandwidth networks for AI training fabrics, inference front-ends, storage, and site/OT networks.
- Design and maintain supercomputer data center and campus networks according to company standards.
- Evaluate, procure, and deploy network hardware, including data-center class switches and related appliances.
- Contribute to maturing network automation tooling and implement configuration analysis, linting, validation, and scalable deployment frameworks.
- Plan and coordinate network change windows with stakeholders.
- Troubleshoot and resolve network-related issues affecting cluster health and job performance.
- Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns.
- Proactively tailor network monitoring and telemetry.
- Continuously create and update network documentation.
- Collaborate with cross-functional teams to identify and resolve potential design issues.
- Perform job walks with customers, vendors, and contractors.
- Ensure networks are configured and maintained in compliance with industry and cybersecurity standards.
Basic Qualifications:
- Bachelor's degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience;
- OR 5+ years of professional network engineering experience in lieu of a degree.
- Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks.
- Functional experience with multiple network vendors in production or lab environments.
- Experience with GitOps and Infrastructure as Code frameworks.
Preferred Skills and Experience:
- Strong understanding of the OSI model and network standards.
- Hands-on experience with Cisco, Arista, Juniper, and/or NVIDIA Spectrum-X data-center class switches.
- Experience with RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience is a plus.
- Working knowledge of AI training and inference traffic patterns.
- Experience with WDM and large-scale single-mode / multimode fiber plants.
- Experience with switch port security, network segmentation, QoS, multicast, and redundancy protocols.
- Familiarity with network monitoring and Layer 1 test tools.
- Proficiency in scripting and automation frameworks.
- Linux and Windows system administration experience.
- Industry-standard certifications such as CCNA or CCNP.
- Experience supporting real-time systems, industrial control / OT networks, or high-reliability environments.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/xai/jobs/5229355007