Description
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.
As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling.
Responsibilities
- Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.
- Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
- Run blameless postmortems and drive corrective actions to closed, not filed.
- Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
- Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
- Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
- Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.
Basic Qualifications
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
- Proven large-scale incident command experience and calm technical leadership on a bridge.
- Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
- Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
- Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
- Excellent problem-solving skills with a data-driven approach to reliability engineering.
- Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
Preferred Skills and Experience
- Experience in AI/ML infrastructure or supercomputing environments.
- Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
- Experience running game days, dependency mapping, and closed-loop corrective action programs.
- Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
- Prior work in a fast-paced startup or tech company like SpaceXAI.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/xai/jobs/5229153007