# Site Reliability Engineer

**Company**: Databricks Information Technology
**Location**: Costa Rica
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/databricks/jobs/8493168002?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_a9baae61-d6b

## Description

As a Site Reliability Engineer at Databricks, you will bridge the gap between software engineering and systems architecture. You will be a core contributor to the IT Infrastructure team, owning the evolution of core infrastructure and observability platforms. This role requires a strong software engineering mindset and deep technical breadth to deliver high-quality, scalable solutions for immature system problems.

Your focus will be on building resilient, automated infrastructure that empowers development teams and ensures our cloud environment is cost-optimised, secure, and highly available.

Key responsibilities include:

- Architect and Automate: Design and deploy production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi.

- Reliability and Performance Engineering: Optimise system performance, architecture, and scaling to ensure maximum uptime and minimal latency for critical IT services.

- CI/CD Excellence: Architect robust deployment pipelines (e.g., GitHub Actions), managing both hosted and self-hosted runners for specialised build requirements.

- Observable by Default: Create underlying infrastructure to ensure new internal applications are secure and have logging, metrics, and alerts enabled by default.

- Agentic Tooling: Build internal AI plugins and automation scripts to streamline developer workflows and enhance operational efficiency.

- Incident Response: Focus on subsequent data usage, incident management workflows, and creating necessary dashboards to maintain service health. Participate in a shared on-call rotation, leading rapid incident response and technical troubleshooting for production outages.

- Partner Cross-Functionally: Collaborate with Security, Engineering, and Support teams to deliver real business outcomes.

## Skills

### Required
- Python
- Infrastructure as Code (IaC)
- Cloud & Containers
- Observability Mindset
- Distributed Systems
- CI/CD Proficiency

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/databricks/jobs/8493168002?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
