Description
As an Incident Manager at Databricks, you will lead critical production incidents, provide clear communication to customers, executives, and engineers, and serve as both incident commander and reliability engineer.
Your work will ensure Databricks maintains technical resilience and customer confidence during high-impact events.
Key responsibilities:
- Lead critical incidents and coordinate multi-disciplinary response efforts
- Drive technical root cause analysis and reliability improvements
- Collaborate with engineering teams to analyze and prevent failures
- Own communications during incidents and deliver frequent updates to stakeholders
- Mentor and train peers in incident communication and technical response disciplines
Requirements:
- 5+ years of experience in incident management, site reliability engineering, or production operations
- Proven ability to lead high-severity incidents and coordinate multi-team response efforts
- Strong understanding of cloud infrastructure (AWS, Azure, or GCP)
- Deep expertise in log analysis and debugging
- Proficiency in at least one major programming or scripting language (Python, Go, or Bash)
- Experience developing and maintaining incident playbooks and communication templates
- Excellent contextual interpretation and writing skills
- BS, Master's or other advanced degree in Computer Science or Computer Engineering
Pay Range: $103,900-$145,525 USD
Benefits: Comprehensive benefits and perks that meet the needs of all employees
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/databricks/jobs/8407869002