Description
We are looking for a Director of Cloud SRE to lead a team of engineering leaders and engineers responsible for federating core SRE principles across a global, hybrid technology organization.
This role drives the strategy, architecture, and roadmap for our internal observability and reliability tooling , spanning telemetry standards, developer pipeline integration, and application team SRE maturity , and works in close partnership with peer SRE leaders to extend that strategy consistently across cloud-native and on-premise environments alike.
The ideal candidate blends technical depth with organizational fluency: someone who can sit in an architecture review and a roadmap planning session with equal credibility, who has personally built and operated production systems, and who can partner effectively across a broader SRE leadership team to advance a long-term, unified observability strategy.
Responsibilities
- Partner with fellow SRE leaders to define and drive a multi-year, holistic strategy for unified observability and SRE platform offerings, spanning cloud-native (GCP) and on-premise (data center, manufacturing, distribution, campus) environments.
- Lead, develop, and grow a team of engineering managers/leads and individual contributor engineers, building organizational depth in SRE practice and platform engineering.
- Drive the roadmap for internally-built observability tooling, ensuring architecture remains vendor-agnostic and portable across telemetry backends (OpenTelemetry-first design, current integration with Dynatrace) with a focus on Agentic AI platforms to simplify correlation data.
- Help federate core SRE principles , SLIs/SLOs, error budgets, incident management, toil reduction, capacity and reliability engineering , across application and platform teams enterprise-wide, working alongside peer SRE leaders rather than centralizing reliability as a bottleneck.
- Partner with developer experience and platform engineering teams to embed observability and reliability tooling directly into CI/CD pipelines and source repositories, shifting reliability left in the development lifecycle.
- Contribute to an SRE maturity model, providing application teams a clear, staged path to deepen their own reliability practice with SRE org support and self-service tooling.
- Build cross-domain relationships with manufacturing, plant, and OT engineering leadership, in partnership with other SRE leaders, to extend reliability and observability discipline into environments with materially different constraints (legacy protocols, air-gapped or constrained networks, safety-critical operations).
- Represent SRE platform direction to senior technology leadership, including architecture governance bodies, and act as an escalation point for major reliability and observability initiatives within your team's scope.
- Contribute to vendor relationship and technology decisions related to observability tooling, balancing build-vs-buy tradeoffs against long-term platform and cost strategy.
- Ensure the team maintains hands-on technical currency , reviewing designs, contributing to architecture decisions, and staying credible as a technical leader.
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
- 10+ years of experience in Site Reliability Engineering, platform engineering, or infrastructure engineering
- 4+ years in a people leadership role managing engineering leaders and/or engineers.
- Demonstrated experience building and operating observability platforms at scale, with hands-on depth in OpenTelemetry and at least one enterprise observability platform (Dynatrace, Datadog, New Relic, Splunk, or similar).
- Proven track record designing and delivering internally-built developer tooling, including integration with CI/CD pipelines, source control platforms, and developer workflows.
- Strong working knowledge of public cloud architecture (GCP strongly preferred; AWS/Azure acceptable) and demonstrated ability to extend reliability practices into hybrid or on-premise environments.
- Deep understanding of core SRE principles: SLIs/SLOs, error budgets, incident management and postmortem practice, toil reduction, capacity planning, and reliability-by-design.
- Experience operating in environments with heterogeneous infrastructure , cloud, data center, and OT/manufacturing or industrial environments a strong plus.
- Demonstrated ability to build a long-term technical strategy and translate it into an executable roadmap, balancing tactical remediation against multi-year platform investment.
- Strong executive communication skills , able to represent technical strategy to senior leadership and align cross-functional stakeholders around a unified direction.
- Experience with infrastructure-as-code (Terraform or equivalent) and modern software delivery practices (agile/PI planning experience a plus).
Preferred:
- Experience building or scaling an SRE function within a large, matrixed enterprise.
- Familiarity with emerging AI/agentic observability standards (OTel GenAI semantic conventions) and their application to platform tooling.
- Experience establishing SRE maturity models or capability frameworks used to guide staged team adoption.
This role requires up to 10% travel.
Benefits
- Immediate medical, dental, vision and prescription drug coverage
- Flexible family care days, paid parental leave, new parent ramp-up programs, subsidized back-up child care and more
- Family building benefits including adoption and surrogacy expense reimbursement, fertility treatments, and more
- Vehicle discount program for employees and family members and management leases
- Tuition assistance
- Established and active employee resource groups
- Paid time off for individual and team community service
- A generous schedule of paid holidays, including the week between Christmas and New Year's Day
- Paid time off and the option to purchase additional vacation time.