Description
GitLab is seeking a Staff Site Reliability Engineer (SRE) to join their Environment Automation team. As an SRE, you will help keep all user-facing services and production systems reliable, scalable, and efficient. The Environment Automation specialization focuses on operating and automating hundreds of GitLab environments, ensuring they remain secure, consistent, and reliable at scale.
Responsibilities:
- Design and implement automation that provisions and manages hundreds of isolated GitLab environments using Terraform, Ansible, and Kubernetes.
- Troubleshoot issues across Kubernetes clusters, cloud services, and GitLab apps.
- Replace manual workflows with infrastructure-as-code solutions.
- Build observability systems that detect bottlenecks, predict usage trends, and optimize resource consumption.
- Lead incident response and postmortem efforts.
- Influence architectural decisions around automation, scalability, and operational excellence.
Requirements:
- Proven ability to operate and troubleshoot production workloads across multiple tenants or environments.
- Strong hands-on experience with Terraform, including workspace strategies, state management, and automation patterns that scale.
- Skilled at diagnosing deployment failures, interpreting pod logs, and debugging scheduling issues and rollback scenarios in live environments.
- Ability to read and debug code in Go and/or Ruby.
- Experience supporting infrastructure for many customers or environments simultaneously.
- Able to reason through complex systems and operational challenges.
Benefits:
- Benefits to support your health, finances, and well-being
- Flexible Paid Time Off
- Team Member Resource Groups
- Equity Compensation & Employee Stock Purchase Plan
- Growth and Development Fund
- Parental Leave
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/gitlab/jobs/8800489002