Description
At Twilio, we're shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences.
As a Software Engineer L2 in Cloud Infrastructure, you will evolve and maintain fundamental Compute infrastructure, collaborating with a passionate team to enhance system capabilities. Key responsibilities include developing scalable cloud-native environments, VM orchestration, AWS-ASG Auto Scaling group of EC2 instances, hardened base AMIs, and secure container images while maintaining critical OS libraries.
Responsibilities:
- Collaborate with Tech Leaders, Architects, and other Engineers to develop solutions for complex problems in distributed computing and infrastructure management.
- Automate solutions for operational issues, such as monitoring, performance, planning, and disaster response.
- Participate in an on-call rotation to support our business-critical infrastructure.
- Ensure high-quality implementation by applying Infrastructure as Code industry standards.
- Demonstrate effective communication by authoring and reviewing design documents, runbooks, and other service documentation, and keeping a good record of changes in the systems.
- Apply Agile methodologies to continuously deliver value to customers.
- Act as a point of contact for legacy/new Compute system components.
Qualifications:
- 2+ years of experience in AWS Cloud infrastructure management (preferably backend/infrastructure-focused like AMI, EC2, IAM policies/roles, etc.).
- Strong ASG (Auto Scaling Groups) knowledge to design, implement, and support scalable cloud-native environments.
- Experience with hardened base AMIs and AL23.
- Proficiency with one or more programming languages, such as Java or Python (includes SW Arch patterns, clean code, debugging, etc.).
- Proficient in shell scripting to streamline repetitive tasks and enhance efficiency in operations.
- Skills to work independently with multiple global teams, developing, configuring, deploying, and operating the global Twilio Infrastructure Platform, blending operational excellence with development best practices.
- Knowledge of container-based application/services.
Desired:
- Knowledge of deployment tools and frameworks like infrastructure as code and continuous deployment processes (e.g., Github, Buildkite, Terraform-TFC, ArgoCD, Harness, Cloud network).
- Operational experience in complex distributed systems, including experience with SLO/SLAs towards high availability and reliability goals, including tools like DataDog or Prometheus.
- Exposure to File Integrity Monitoring (FIM) tools, specifically Falco, and awareness of compliance frameworks (PCI, SOX).
- Knowledge of Kubernetes.
- Experience with Claude AI or similar.
Location: This role will be remote, but is not eligible to be hired in San Francisco, CA, Oakland, CA, San Jose, CA, or the surrounding areas.
Travel: You may be required to travel occasionally to participate in project or team in-person meetings.
What We Offer: Working at Twilio offers many benefits, including competitive pay, generous time off, ample parental and wellness leave, healthcare, a retirement savings program, and much more. Offerings vary by location.
Compensation: The estimated pay ranges for this role are as follows:
- Based in Colorado, Hawaii, Illinois, Maryland, Massachusetts, Minnesota, Vermont, or Washington D.C.: $116,960.00-$146,200.00
- Based in New York, New Jersey, Washington State, or California (outside of the San Francisco Bay area): $123,760.00-$154,700.00
- Based in the San Francisco Bay area, California: $137,520.00-$171,900.00
- This role may be eligible to participate in Twilio's equity plan and corporate bonus plan.