# Senior Staff Infrastructure Engineer

**Company**: VGS
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://jobs.lever.co/verygoodsecurity/5dd25169-ee77-4e17-bb07-ca77ebb04359?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_237986ca-b47

## Description

As a Senior Staff Infrastructure Engineer, you will serve as a technical leader on our Platform Engineering team. You will architect, scale, and fortify global cloud infrastructure designed to handle mission-critical, high-throughput payments applications with zero downtime.

You will take ownership of key platform foundations powering our core payments infrastructure. Rather than executing against a rigid task list or holding blanket ownership over the entire platform, you will drive the technical strategy, architecture, and reliability standards for your designated domains within our multi-region AWS environment.

## Responsibilities

- Build Immutable, Self-Healing Systems: Design, build, and optimize multi-region, high-availability AWS infrastructure. You will drive our evolution from hand-crafted environments to a standardized, globally scalable fleet managed entirely through code.

- Drive Resiliency & Automation: Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and IaC.

- Deep Observability & Resiliency: Build end-to-end telemetry (Prometheus, Grafana, OpenTelemetry) to proactively spot bottlenecks. You will own incident management and conduct blameless post-mortems to continuously harden our reliability baseline.

- Force-Multiply Engineering Velocity: Partner closely with Product, Security, and Core Engineering teams and lead from the front by designing "Golden Paths" that strip away friction for feature teams. You will influence company-wide engineering practices and mentor the organization on how to move fast with high alignment.

- Customer Impact: Architect and operate high-performance, low-latency private connectivity to optimize the experience for external customers. Partner strategically with internal engineering teams at the design and architectural level for platform enablement and adoption.

## Requirements

- Ownership at Scale: 10+ years of experience taking personal ownership of outcomes in complex, large-scale distributed systems within mission-critical environments.

- AWS & Infrastructure-as-Code: Advanced proficiency in AWS ecosystems leveraging Terraform to build reproducible environments.

- Containerization & Orchestration: Strong, hands-on experience with Kubernetes (EKS), Docker, and GitOps workflows (Flux, Argo, GitHub Actions).

- Automation & Scripting: Strong coding skills in Python, Go, or Bash to automate infrastructure and build operational tools.

- Observability Expertise: Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.

- Security & Networking Foundations: Solid understanding of cloud security, API Gateways, load balancing, and network isolation, viewing security as a fundamental engineering constraint, not an afterthought.

## Skills

### Required
- AWS
- Terraform
- Kubernetes
- Docker
- GitOps
- Python
- Go
- Bash
- Prometheus
- Grafana
- OpenTelemetry

### Nice to have
- tokenization
- payment processing
- cryptology
- security products
- Kafka
- database performance tuning
- Java
- Spring Framework

---

Source: [Apply at jobs.lever.co](https://jobs.lever.co/verygoodsecurity/5dd25169-ee77-4e17-bb07-ca77ebb04359?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
