Description
SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.
The Data Center Engineering team builds internal systems and platforms that keep data centers running at the scale and reliability required for frontier AI training and inference. The team partners closely with datacenter operations, research, and infrastructure teams to deliver high-leverage tools that turn raw operational data into clear insight and action.
Responsibilities:
- Build and operate software stacks for site operations, including repair trackers, vendor turnback workflows, operational dashboards, and their integrations.
- Design, build, and operate multi-service production systems for systems like SRT-style repair/maintenance trackers and vendor turnover/turnback state machines.
- Ensure correctness of operational state, including node state accuracy, queue ownership, and audit trails.
- Build and maintain integrations with ticketing, inventory/rack systems, telemetry stores, and vendor portals.
- Keep tools reliable, focusing on uptime, data integrity, reconciliation, access control, and safe deploys.
- Embed with SiteOps and NOC users, measuring workflow adoption.
Basic Qualifications:
- Bachelor's degree in Computer Science, Engineering, or related fields.
- 3+ years of experience building and operating production software.
- Strong fundamental knowledge of computer science, including data structures, algorithms, operating systems, and networking.
- Strong proficiency in at least one programming language.
- Experience designing and developing RESTful APIs.
- Experience working with databases like Postgres, MongoDB, MySQL, DynamoDB, etc.
- Experience collaborating with cross-functional teams.
- Experience working in cloud platforms like GCP, Microsoft Azure, AWS, OCI, or similar.
- Experience using observability tools and dashboards.
- Experience writing unit tests and integration tests.
Preferred Skills and Experience:
- Full-stack or backend + data experience, especially with workflow and state-machine systems.
- Experience automating deployments using CI/CD tools.
- Experience building high-correctness operational UIs.
- On-call discipline and a track record of treating internal platforms with production rigor.
- Experience designing and shipping multi-service systems.
- Experience delivering high-quality internal tools in rapidly changing environments.
- Proven ownership of correctness-sensitive systems.
- Experience integrating with external systems via APIs.
- MS in Computer Science or related field.
- Experience collaborating closely with operations, NOC, and infrastructure teams.
- Experience working with performance load testing tools.
Additional Requirements:
- The role is fully onsite in Memphis, TN or Southhaven, MS.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/xai/jobs/5209858007