# Senior Customer Reliability Engineer

**Company**: Cloudflare
**Location**: Singapore
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://job-boards.greenhouse.io/cloudflare/jobs/8037611?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_4ab54716-8ef

## Description

## Job Description

At Cloudflare, we're on a mission to help build a better Internet. Today, the company runs one of the world's largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies.

## Why This Role Exists

Cloudflare built its reputation helping build a better Internet, defending millions of sites, giving away SSL and DDoS mitigation when the industry charged premium prices. In an acceleratingly dangerous world, the scope of that mission has changed. We are becoming something more: critical infrastructure. Banks run their payment rails on us. Governments run public services on us. Media companies depend on us during live events. Health systems depend on us to provide care. Reliability for these customers is no longer a feature of our product. It is a mission.

## The Role

The Customer Reliability Engineering function is the spine of that pivot. CRE is SRE applied outward, the same engineering discipline, applied to the reliability of the systems our customers run on Cloudflare. You are the engineer who owns the problems that matter most to the customers who matter most, and you contribute directly to our products and tooling, in partnership with Product Engineering, to hold that standard across the entire customer base.

## Responsibilities

- Rapid incident response and root cause analysis. Own the most complex, high-severity customer issues end-to-end, from first signal through confirmed resolution.

- Lead deep-dive debugging across the full stack: edge, network, DNS, transport, APIs, application, customer-side configuration.

- Reproduce defects, validate fixes with Engineering, and confirm customer-side resolution.

- Produce postmortems other engineers rely on.

- Hold on-call for high-severity incidents as part of a global rotation that includes weekends.

- Proactive reliability engineering. Analyze support and telemetry signals across the customer base to find systemic risks before they become incidents.

- Contribute monitoring, detection, and diagnostic capability to the core product and the engineering systems that give Customer Support early visibility into customer-affecting issues.

- Define customer-facing reliability metrics (error rates, resolution times, repeat-contact rates) and drive measurable improvement.

- Write automation that reduces mean-time-to-detect and mean-time-to-resolve.

- Cross-functional partnership. Manage the technical escalation lifecycle with clear ownership and timely communication.

- Partner with Product Engineering to drive fixes, workarounds, and configuration changes that address underlying gaps.

- Represent the customer reliability perspective in engineering syncs, incident reviews, and post-mortem processes.

- Technical leadership and enablement. Raise the technical floor of Customer Support through pair-debugging, structured knowledge transfer, and shared tooling.

- Document diagnostic procedures and resolution patterns in runbooks, internal knowledge bases, and AI skills.

- Share insights from customer-facing incidents to improve product documentation and operational readiness.

## Requirements

- Minimum 5 years of hands-on experience in site reliability engineering, escalation engineering, systems engineering, or a comparable deeply technical support / operations role, with at least 2 years in customer-facing environments.

- Strong foundation in networking and security: TCP/IP fundamentals, core protocols, routing protocols, firewall concepts, VPN and encryption, Zero Trust architecture.

- Proficiency with observability and diagnostic tooling: packet capture and analysis, log aggregation, metrics dashboards, distributed tracing.

- Strong scripting and automation skills (Bash, Python) with a track record of shipping tooling that improves reliability and reduces toil.

- Experience with incident management, postmortem culture, and SLO/SLI-based reliability practices.

- Excellent written and verbal communication.

## Desired Skills & Experience

- SRE, DevOps, or platform engineering experience with direct customer-facing accountability.

- Deep expertise at both L3/L4 (network infrastructure) and L7 (application protocols, DNS, HTTP, WebSocket).

- Expert-level proficiency with Linux command-line tools.

- Data-at-scale analysis using SQL, PromQL, or equivalent.

- Familiarity with CI/CD pipelines, infrastructure-as-code, and container orchestration.

- Track record of building internal tooling or diagnostic utilities that measurably improved team efficiency.

- Demonstrated technical leadership: mentoring engineers, driving cross-team initiatives, influencing outcomes without direct authority.

- Experience applying AI/ML to production engineering or operational workflows.

## What Makes Cloudflare Special?

We're not just a highly ambitious, large-scale technology company. We're a highly ambitious, large-scale technology company with a soul. Fundamental to our mission to help build a better Internet is protecting the free and open Internet.

## Skills

### Required
- TCP/IP fundamentals
- networking and security
- observability and diagnostic tooling
- scripting and automation skills
- incident management
- postmortem culture
- SLO/SLI-based reliability practices
- written and verbal communication

### Nice to have
- SRE
- DevOps
- platform engineering
- L3/L4 expertise
- L7 expertise
- Linux command-line tools
- data-at-scale analysis
- CI/CD pipelines
- infrastructure-as-code
- container orchestration
- technical leadership
- AI/ML

---

Source: [Apply at job-boards.greenhouse.io](https://job-boards.greenhouse.io/cloudflare/jobs/8037611?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
