Description
As a Senior Staff Infrastructure Engineer, you will serve as a technical leader on our Platform Engineering team. You will architect, scale, and fortify global cloud infrastructure designed to handle mission-critical, high-throughput payments applications with zero downtime.
You will take ownership of key platform foundations powering our core payments infrastructure. Rather than executing against a rigid task list or holding blanket ownership over the entire platform, you will drive the technical strategy, architecture, and reliability standards for your designated domains within our multi-region AWS environment.
Responsibilities
- Build Immutable, Self-Healing Systems: Design, build, and optimize multi-region, high-availability AWS infrastructure. You will drive our evolution from hand-crafted environments to a standardized, globally scalable fleet managed entirely through code.
- Drive Resiliency & Automation: Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and IaC.
- Deep Observability & Resiliency: Build end-to-end telemetry (Prometheus, Grafana, OpenTelemetry) to proactively spot bottlenecks. You will own incident management and conduct blameless post-mortems to continuously harden our reliability baseline.
- Force-Multiply Engineering Velocity: Partner closely with Product, Security, and Core Engineering teams and lead from the front by designing "Golden Paths" that strip away friction for feature teams. You will influence company-wide engineering practices and mentor the organization on how to move fast with high alignment.
- Customer Impact: Architect and operate high-performance, low-latency private connectivity to optimize the experience for external customers. Partner strategically with internal engineering teams at the design and architectural level for platform enablement and adoption.
Requirements
- Ownership at Scale: 10+ years of experience taking personal ownership of outcomes in complex, large-scale distributed systems within mission-critical environments.
- AWS & Infrastructure-as-Code: Advanced proficiency in AWS ecosystems leveraging Terraform to build reproducible environments.
- Containerization & Orchestration: Strong, hands-on experience with Kubernetes (EKS), Docker, and GitOps workflows (Flux, Argo, GitHub Actions).
- Automation & Scripting: Strong coding skills in Python, Go, or Bash to automate infrastructure and build operational tools.
- Observability Expertise: Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.
- Security & Networking Foundations: Solid understanding of cloud security, API Gateways, load balancing, and network isolation, viewing security as a fundamental engineering constraint, not an afterthought.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://jobs.lever.co/verygoodsecurity/5dd25169-ee77-4e17-bb07-ca77ebb04359