# Manager, Operations

**Company**: xAI
**Location**: Memphis, TN
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: Operations
**Industry**: Technology
**Wikidata**: https://www.wikidata.org/wiki/Q120599684

**Apply**: https://job-boards.greenhouse.io/xai/jobs/5125589007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_106d8bed-bea

## Description

## About the Role

We are seeking an exceptional Manager, Operations to lead facilities operations and power generation for xAI's hyperscale AI compute facilities. This role will own the day-to-day and long-term performance of mission-critical data center operations, including power generation, power distribution, cooling, mechanical, electrical, and environmental systems, while also directing the fiber teams responsible for high-capacity networking and connectivity that support our supercomputing clusters.

You will build and lead high-performing operations, power generation, and fiber teams, drive relentless reliability and efficiency, and ensure seamless 24/7 uptime for the infrastructure powering xAI's AI training at unprecedented scale. This high-impact position requires deep expertise in data center or hyperscale operations (including power generation), strong leadership in fast-paced environments, and the ability to deliver world-class performance under aggressive growth timelines.

## Responsibilities

- Lead and scale the facilities operations and power generation teams responsible for the reliable operation, maintenance, monitoring, and optimization of critical infrastructure including on-site power generation assets, electrical systems, mechanical/HVAC, liquid cooling, power distribution, UPS, generators, and building management systems.

- Direct the fiber teams overseeing the design, deployment, maintenance, and expansion of high-speed fiber optic networks, dark fiber, and connectivity infrastructure supporting AI compute clusters and data center interconnects.

- Own key performance metrics such as uptime (targeting 99.999%+), mean time to detect/repair (MTTD/MTTR), power usage effectiveness (PUE), water usage effectiveness (WUE), power generation efficiency, and overall infrastructure availability.

- Develop and enforce standard operating procedures (SOPs), preventive maintenance programs, incident response protocols, and continuous improvement processes for both facilities and power generation assets to minimize downtime and maximize efficiency.

- Build, mentor, and grow multidisciplinary teams of operations technicians, power generation engineers and controls specialists while fostering a culture of ownership, safety, and excellence.

- Partner closely with engineering, construction, procurement, and AI hardware teams to support new facility builds, expansions, commissioning, power integration, and smooth handovers from project to operations.

- Manage operational budgets, vendor relationships (maintenance contractors, fiber providers, power generation OEMs, fuel suppliers), spare parts inventory, and risk mitigation strategies in a high-velocity environment.

- Drive innovation in operational practices, automation, predictive maintenance, power generation optimization, and sustainability initiatives to support the extreme power and cooling demands of next-generation AI systems.

- Provide regular performance reporting, root cause analyses, lessons learned, and strategic recommendations to senior leadership.

## Basic Qualifications

- 5+ years of progressive experience in data center facilities operations, power generation operations, hyperscale infrastructure management, or mission-critical industrial operations, with at least 2+ years in a management or supervisor role.

- Proven track record leading large-scale operations teams supporting high-density compute environments with significant on-site or dedicated power generation (AI, HPC, or hyperscaler data centers strongly preferred).

- Strong experience managing fiber optic networks, dark fiber deployments, or high-bandwidth connectivity infrastructure in large-scale technical environments.

- Deep knowledge of power generation systems (gas turbines, reciprocating engines, cogeneration, etc.), MEP (mechanical, electrical, plumbing) systems, BMS/SCADA, liquid cooling, power redundancy topologies, and 24/7 operations best practices.

- Demonstrated success delivering high reliability, rapid incident resolution, and operational excellence under aggressive scaling timelines.

- Hands-on leadership style with the ability to roll up sleeves while effectively managing teams, budgets, and cross-functional stakeholders.

- Proficiency with operations tools, CMMS (computerized maintenance management systems), monitoring platforms, and data-driven decision making.

## Skills

### Required
- data center operations
- power generation
- fiber optic networks
- leadership
- operations management

### Nice to have
- AI or hyperscale data center operations
- liquid cooling systems
- high-power GPU/accelerator environments
- on-site power generation
- automation
- predictive analytics

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/xai/jobs/5125589007?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
