Description
GitLab is seeking a Backend Engineer for its Geo team. As a Backend Engineer on the Geo team, you'll design, build, and maintain the backend services that power GitLab Geo, Disaster Recovery, and Backup and Restore.
Working primarily in Ruby on Rails and PostgreSQL, you'll improve replication, verification, failover, and data migration for large self-managed and GitLab Dedicated environments. You'll collaborate with Database Engineering, Infrastructure, GitLab Dedicated, and other teams to help distributed customers access repositories faster and recover reliably, while maintaining a high bar for observability, careful rollout, and operational resilience.
Key responsibilities:
- Design, build, and maintain backend functionality for Geo, Disaster Recovery, and Backup and Restore using Ruby on Rails and related services.
- Improve replication and verification workflows for repositories and related data, with a focus on performance, correctness, and operational simplicity.
- Use PostgreSQL features, including logical replication, to support Geo, GitLab Dedicated migrations, and future Cells architectures.
- Measure and reduce replication lag and verification failures through logs, metrics, and alerts, identifying disaster recovery risks early.
- Collaborate with Database Engineering, Infrastructure, GitLab Dedicated, Tenant Scale, and other teams on initiatives that affect Geo and migrations.
- Participate in incident response and post-incident reviews, then turn findings into product and operational improvements.
- Partner with Support, Site Reliability Engineering, and Customer Success on Requests for Help, customer escalations, complex migrations, and incidents where Geo capabilities are critical.
- Own projects from proposal and design through implementation, review, rollout, and production monitoring, while providing constructive feedback on merge requests.
Requirements:
- Experience building and maintaining Ruby on Rails applications in production environments.
- Experience with PostgreSQL or similar relational databases, including replication, indexing, and performance tuning.
- Understanding of distributed systems and data replication concepts, including consistency, eventual consistency, and failure modes.
- Experience building and operating background job or worker systems, such as Sidekiq, including monitoring and retry strategies.
- Knowledge of replication and verification patterns, such as ordered event delivery and checksums, and experience debugging issues such as replication lag.
- Experience operating or debugging large production deployments across installation methods, operating systems, or cloud providers.
- Familiarity with recovery point objective (RPO), recovery time objective (RTO), planned failover, and backup and restore for stateful systems.
Benefits:
- Benefits to support your health, finances, and well-being
- Flexible Paid Time Off
- Team Member Resource Groups
- Equity Compensation & Employee Stock Purchase Plan
- Growth and Development Fund
- Parental Leave
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/gitlab/jobs/8695515002