New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
VGS

Senior Staff Infrastructure Engineer

VGS
Apply →
senior full-time

First indexed 2 Sept 2026

Description

As a Senior Staff Infrastructure Engineer, you will serve as a technical leader on our Platform Engineering team. You will architect, scale, and fortify global cloud infrastructure designed to handle mission-critical, high-throughput payments applications with zero downtime.

You will take ownership of key platform foundations powering our core payments infrastructure. Rather than executing against a rigid task list or holding blanket ownership over the entire platform, you will drive the technical strategy, architecture, and reliability standards for your designated domains within our multi-region AWS environment.

Responsibilities

  • Build Immutable, Self-Healing Systems: Design, build, and optimize multi-region, high-availability AWS infrastructure. You will drive our evolution from hand-crafted environments to a standardized, globally scalable fleet managed entirely through code.
  • Drive Resiliency & Automation: Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and IaC.
  • Deep Observability & Resiliency: Build end-to-end telemetry (Prometheus, Grafana, OpenTelemetry) to proactively spot bottlenecks. You will own incident management and conduct blameless post-mortems to continuously harden our reliability baseline.
  • Force-Multiply Engineering Velocity: Partner closely with Product, Security, and Core Engineering teams and lead from the front by designing "Golden Paths" that strip away friction for feature teams. You will influence company-wide engineering practices and mentor the organization on how to move fast with high alignment.
  • Customer Impact: Architect and operate high-performance, low-latency private connectivity to optimize the experience for external customers. Partner strategically with internal engineering teams at the design and architectural level for platform enablement and adoption.

Requirements

  • Ownership at Scale: 10+ years of experience taking personal ownership of outcomes in complex, large-scale distributed systems within mission-critical environments.
  • AWS & Infrastructure-as-Code: Advanced proficiency in AWS ecosystems leveraging Terraform to build reproducible environments.
  • Containerization & Orchestration: Strong, hands-on experience with Kubernetes (EKS), Docker, and GitOps workflows (Flux, Argo, GitHub Actions).
  • Automation & Scripting: Strong coding skills in Python, Go, or Bash to automate infrastructure and build operational tools.
  • Observability Expertise: Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.
  • Security & Networking Foundations: Solid understanding of cloud security, API Gateways, load balancing, and network isolation, viewing security as a fundamental engineering constraint, not an afterthought.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://jobs.lever.co/verygoodsecurity/5dd25169-ee77-4e17-bb07-ca77ebb04359