Description
At Cloudflare, we're on a mission to help build a better Internet. The company runs one of the world's largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code.
The Product Platform Tools team builds and operates Cloudflare's internal support and admin platform - the foundation that customer-facing and operational teams across the company rely on to do their jobs. As a Systems Engineer on this team, you'll design and build the backend services, infrastructure, APIs, and integrations that power this platform at enterprise scale, working closely with both engineering teams and non-engineering stakeholders like Support, Legal, and Security.
Responsibilities:
- Design, implement, and operate backend services, infrastructure, APIs, and integrations that form the core of Cloudflare's internal support and admin platform, built to scale with the demands of a large, globally distributed organization.
- Partner with engineering teams across the company to onboard their support and admin applications onto the platform - understanding their requirements and ensuring smooth, well-governed, and scalable integration.
- Work directly with non-engineering stakeholders - including Support, Legal, and Security teams - to understand their operational needs and translate them into reliable, maintainable, and appropriately scaled platform capabilities.
- Own and continuously improve the platform's security and authorization posture, partnering with the Security and Legal team to ensure access control policies, governance requirements, and compliance standards are correctly implemented and enforced at scale.
- Ramp quickly into unfamiliar and legacy codebases, build context across a wide range of Cloudflare systems, and make thoughtful improvements that reduce operational toil and improve platform reliability.
- Participate in on-call rotations, respond to production incidents, and drive post-incident reviews to raise reliability and scalability standards across the platform.
- Contribute to platform observability, developer experience, and engineering standards with broad impact across the teams that depend on this platform.
Must-Have Skills:
- 5+ years of professional experience building and operating production backend services at enterprise or large-scale organizations, with strong proficiency in Go and/or TypeScript.
- Demonstrated experience in platform engineering, internal tooling, or support/admin systems , with direct experience supporting non-technical end users or customer-facing teams.
- Strong understanding of authorization and authentication systems such as OAuth 2.0, OIDC, JWT, OPA/Rego, role-base access control, policy based-access or similar and a security-first approach to system design.
- Proven scalability mindset , experience designing and operating systems that serve a large, diverse set of internal consumers with varying reliability and performance requirements.
- Proven ability to work across organizational boundaries: comfortable engaging with engineering teams, Security, Legal, and non-technical operational stakeholders alike.
- Confident working in unfamiliar and legacy codebases; able to ramp quickly, orient independently, and contribute incrementally without full context.
- Strong written and verbal communication , able to translate between technical implementation details and the business or operational needs of diverse stakeholders.
- Advanced, hands-on experience with Kubernetes and Docker in production environments , including cluster operations, resource management, networking, and debugging containerized workloads.
- Strong database experience , including schema design, query optimization, and operating databases in high-availability production environments.
Nice-to-Have Skills:
- Experience building or operating internal support tooling, admin consoles, case management systems, or operational dashboards used by non-engineering teams at enterprise scale.
- Prior exposure to security governance workflows, compliance tooling, or partnering closely with a dedicated Security or Legal team on access control requirements.
- Familiarity with Cloudflare's product ecosystem (Workers, Access, Zero Trust, Gateway) as either a user or implementer.
- Experience with observability tooling and practices , including distributed tracing, structured logging, and metrics dashboards using tools such as Grafana, Kibana, Jaeger, or similar.
- Experience with Terraform for infrastructure-as-code and GitLab CI/CD for automated build, test, and deployment pipelines.
- Exposure to Temporal workflows or distributed workflow orchestration.