Description
At Scale AI, our mission is to accelerate the development of AI applications.
We are looking for an Infrastructure Engineering Manager to help shape the future of AI-powered applications. In this role, you'll bridge the gap between AI research and production, leading a team that is turning innovative prototypes into scalable, high-performance enterprise solutions.
Responsibilities:
- Define and execute the infrastructure roadmap aligned with business and engineering priorities.
- Lead the design and implementation of scalable, secure, and reliable infrastructure systems.
- Set and maintain SLAs/SLOs for platform uptime, performance, and developer experience.
- Manage the engineering team and drive technical delivery.
- Design, build, and optimize backend services for advanced AI-driven applications, focusing on AI agents, evaluation tooling, and automation.
- Influence the culture, values, and processes of a growing engineering team.
- Inspire and mentor engineers.
- Work closely with product, security, and engineering leadership to align on goals and priorities.
Requirements:
- At least 5 years of relevant experience and at least 2+ years of experience managing infrastructure or platform teams.
- Proven experience with cloud platforms such as AWS, GCP, Azure or OCI.
- Proven experience with kubernetes deployments on on-prem infrastructure.
- Deep understanding of CI/CD pipelines, infrastructure-as-code, and container orchestration.
- Experience managing production environments with high availability, reliability, and scalability requirements.
- Familiarity with monitoring, alerting, and incident response best practices.
- Experience working with modern developer platforms and internal tooling to improve engineering velocity.
- Solid foundation and real-world experience in network engineering.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/scaleai/jobs/4719479005