Description
Job Overview
We are seeking a highly skilled Senior Database Reliability Engineer (DBRE) with expertise in PostgreSQL at scale and experience with MySQL. As a DBRE, you will design, operationalize, and optimise the data persistence layer that powers our large-scale, mission-critical systems.
Responsibilities
Architecture, Reliability & Performance
- Design, implement, and operate highly available PostgreSQL clusters, including physical replication, logical replication, sharding/partitioning, and failover automation.
- Optimise query performance, indexing strategies, schema design, and storage engines.
- Perform capacity planning, growth forecasting, and workload modelling.
- Develop and implement high-availability strategies, including automatic failover, multi-AZ/multi-region setups, and disaster recovery.
Automation & Tooling
- Develop automation for tasks such as provisioning, configuration, backups, failovers, vacuum tuning, and schema management using tools like Terraform, Ansible, Kubernetes Operators, or custom tooling.
- Build monitoring, alerting, and self-healing systems for PostgreSQL and MySQL.
Operations & Incident Response
- Lead response during database incidents, including performance regressions, replication lag, deadlocks, bloat issues, and storage failures.
- Conduct root-cause analysis and implement permanent fixes.
Cross-Functional Collaboration
- Collaborate with software engineers to review SQL, optimise schemas, and ensure efficient use of PostgreSQL features.
- Provide guidance on database-related design patterns, migrations, version upgrades, and best practices.
Required Qualifications
- 4+ years of hands-on PostgreSQL experience in high-volume, distributed, or large-scale production environments.
- Strong knowledge of PostgreSQL internals, including WAL, MVCC, bloat/vacuum tuning, query planner, indexing, and logical replication.
- Production experience with MySQL, including InnoDB internals, replication, and performance tuning.
- Advanced SQL skills and strong understanding of schema design and query optimisation.
- Experience with Linux systems, networking fundamentals, and systems troubleshooting.
- Experience building automation with Go or Python.
- Production experience with monitoring tools like Prometheus, Grafana, Datadog, PMM, and pg_stat_statements.
- Hands-on experience with cloud environments, such as AWS or GCP.
Preferred/Bonus Qualifications
- Experience with PgBouncer, HAProxy, or other connection-pooling/load-balancing layers.
- Exposure to event streaming (Kafka, Debezium) and change data capture.
- Experience supporting 24/7 production environments with on-call rotation.
- Contributions to open-source PostgreSQL ecosystem.
What We Offer
- Competitive salary: $160,000 - $220,000 USD per year (for San Francisco Bay Area)
- Equity (where applicable)
- Bonus
- Health, dental, and vision insurance
- 401(k)
- Flexible spending account
- Paid leave (including PTO and parental leave)
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/okta/jobs/7617976