Description
GitLab is looking for an Intermediate Site Reliability Engineer to join the Runners Platform team. You will build and operate Hosted Runners for GitLab Dedicated, the managed CI/CD compute platform that runs customers' pipelines inside single-tenant Amazon Web Services (AWS) environments.
Responsibilities:
- Design, build, and operate AWS infrastructure for Hosted Runners across many single-tenant environments.
- Develop and maintain infrastructure as code using Terraform.
- Write Go code for runner tooling and autoscaling components.
- Build and improve the GitLab CI/CD pipelines that orchestrate blue/green zero-downtime deployments.
- Define and monitor service level objectives for CI job execution.
- Participate in an on-call rotation and handle incidents affecting customer CI/CD workloads.
- Run performance and scale testing that reflects real customer workloads.
- Write documentation and runbooks so the broader team can operate runner stacks consistently.
Requirements:
- Professional experience operating production infrastructure on AWS at scale.
- Strong infrastructure-as-code experience with Terraform.
- Proficiency in Go for building and debugging infrastructure tooling.
- Practical knowledge of CI/CD systems and job execution.
- Experience with observability practices such as metrics, dashboards, alerting, logging.
- Experience with on-call rotations and incident management for customer-facing systems.
- Strong problem-solving skills, excellent written communication.
Benefits:
- Benefits to support your health, finances, and well-being.
- Flexible Paid Time Off.
- Team Member Resource Groups.
- Equity Compensation & Employee Stock Purchase Plan.
- Growth and Development Fund.
- Parental Leave.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/gitlab/jobs/8644274002