Description
We are seeking a Staff Software Engineer to join the Lakeflow Disaster Recovery team at Databricks. The successful candidate will design and implement distributed systems that replicate and recover pipelines across regions.
The Lakehouse is a unified platform for data engineering, analytics, and AI, addressing major challenges in enterprise data platforms. Lakeflow is a critical part of this vision, enabling customers to build and operate streaming and batch ETL pipelines.
As a Staff Engineer on the Lakeflow Disaster Recovery team, you will work on designing and implementing distributed systems that replicate and recover pipelines across regions. This involves solving complex problems related to consistency, idempotency, causal ordering, failover, failback, and safe recovery without silent data loss or duplication.
Key areas of focus include:
- Cross-region replication and recovery for Lakeflow pipelines, streaming tables, and materialized views
- Distributed consistency across pipeline dependencies, table versions, and transaction logs
- Failover and failback workflows with conservative correctness guardrails
- Deep clone, metadata reconciliation, observability, and failure-injection testing
- High-fidelity recovery simulations, game-day testing, and formal reasoning about failure modes
Requirements:
- Passion for distributed systems, databases, storage systems, streaming systems, or reliability engineering
- Strong software engineering skills in Java, Scala, C++, Go, Python, or similar production languages
- Understanding of consistency, transactions, idempotency, replication, checkpointing, and data lineage
- Ability to define and work towards a multi-year technical vision with incremental, production-quality deliverables
- 8+ years of experience working on related systems preferred
- Optional: PhD or advanced research experience in databases, distributed systems, or storage
This role offers the opportunity to build the foundations of resilient Lakehouse computing and help customers keep their data pipelines running when failure matters most.
The pay range for this role is $192,000-$260,000 USD.