# Site Reliability Engineer

**Company**: SpaceXAI
**Location**: Memphis, TN
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/xai/jobs/5229153007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_30b4237f-530

## Description

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling.

## Responsibilities

- Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.

- Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.

- Run blameless postmortems and drive corrective actions to closed, not filed.

- Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.

- Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).

- Define error budgets and availability objectives at campus and service boundaries as adopted by the business.

- Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.

## Basic Qualifications

- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).

- 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.

- Proven large-scale incident command experience and calm technical leadership on a bridge.

- Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.

- Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.

- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.

- Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.

- Excellent problem-solving skills with a data-driven approach to reliability engineering.

- Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.

## Preferred Skills and Experience

- Experience in AI/ML infrastructure or supercomputing environments.

- Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.

- Experience running game days, dependency mapping, and closed-loop corrective action programs.

- Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.

- Prior work in a fast-paced startup or tech company like SpaceXAI.

## Skills

### Required
- site reliability
- systems engineering
- large-scale production operations
- incident command
- monitoring and observability
- scripting
- problem-solving

### Nice to have
- AI/ML infrastructure
- supercomputing environments
- SLOs
- SLIs
- error budgets
- game days
- dependency mapping

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/xai/jobs/5229153007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
