New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
SpaceXAI

Network Operations Center Specialist - Memphis

SpaceXAI
Apply →
onsite senior full-time Southaven, MS

First indexed 4 Sept 2026

Description

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.

As a Network Operations Center (NOC) Specialist, you are the eyes and the voice of the campus , never the hands. You watch campus health signals around the clock, detect and verify site-impacting events, assemble the right responders fast, and run incident communications leadership can trust. You make sure no major incident closes without a timeline, a report, and a tracked corrective project.

Responsibilities:

  • Monitor campus health signals 24/7, detecting and verifying site-impacting events.
  • Classify and log incident dispositions; feed noise patterns back to SRE for signal quality improvement.
  • Operate the escalation matrix and communicate effectively with stakeholders.
  • Run incident bridges and maintain incident timelines.
  • Produce initial root cause analysis (RCA) framing for further investigation by SRE/Hardware Failure Analysis.
  • Write major-incident reports and drive corrective projects to closure.
  • Continuously improve NOC processes, runbooks, and communications templates.

Basic Qualifications:

  • Experience in a 24/7 operations environment (NOC, SOC, dispatch, mission control, or equivalent).
  • Proven ability to handle incidents under SLA in a high-signal environment.
  • Experience with incident bridges, stakeholder updates, and timeline management.
  • Excellent written and verbal communication skills.
  • Pattern recognition across multiple domains (compute, network, storage, facilities).
  • Experience with operational processes (runbooks, escalation matrices).
  • Ability to work rotating shifts, including nights and weekends.

Preferred Skills and Experience:

  • Prior NOC, data center operations, or campus reliability experience in high-performance computing or AI/ML infrastructure.
  • Experience writing major-incident reports and driving corrective actions.
  • Familiarity with Linear or similar work-tracking tools.
  • Experience partnering with SRE, SiteOps, and Facilities on escalations.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/xai/jobs/5229807007