Description
Shield AI is seeking a Senior Cloud Platform Engineer to design, build, and operate the shared technology platforms that support its engineering, autonomy development, and enterprise operations environments.
As part of the Cloud & Infrastructure team within Enterprise Operations, this role will develop standardized infrastructure, automation, and developer-facing capabilities across Shield AI's public and private cloud environments.
Responsibilities
- Design, build, and operate scalable platform infrastructure across Azure, AWS, and private cloud environments.
- Develop reusable infrastructure-as-code modules, automation frameworks, and self-service platform capabilities for engineering teams.
- Own platform initiatives from technical design and implementation through production operations and continuous improvement.
- Build and maintain container and Kubernetes platform capabilities, including deployment automation, configuration management, observability, and lifecycle management.
- Develop platform tooling and automation using technologies such as Terraform, Ansible, Python, and Go.
- Build standardized paved paths that simplify infrastructure provisioning, application deployment, and access to shared platform services.
- Establish and maintain platform standards for reliability, scalability, security, performance, and cost efficiency.
- Support capacity planning, performance tuning, vulnerability remediation, disaster recovery, and platform lifecycle management.
- Build and improve CI/CD pipelines, monitoring, logging, alerting, and other developer-facing platform services.
- Troubleshoot complex platform, infrastructure, and application integration issues and lead root-cause analysis.
- Develop and maintain architecture diagrams, technical documentation, operational procedures, and reusable implementation patterns.
- Evaluate emerging technologies and recommend solutions that improve automation, reliability, and developer productivity.
- Provide technical guidance to other engineers and contribute to platform architecture, standards, and roadmap decisions.
- Participate in an on-call rotation and scheduled after-hours maintenance as needed.
Requirements
- 7+ years of experience in platform engineering, cloud infrastructure, DevOps, Site Reliability Engineering, or a related discipline.
- Experience designing and operating production infrastructure in Azure, AWS, or comparable public cloud environments.
- Strong experience developing reusable infrastructure-as-code modules and automation using tools such as Terraform and Ansible.
- Hands-on experience deploying and operating containerized workloads and Kubernetes-based platforms.
- Proficiency in Python, Go, or another programming language used to build platform tooling and automation.
- Experience building CI/CD pipelines, self-service infrastructure capabilities, or developer-facing platform services.
- Strong Linux systems administration experience, including deployment, configuration, troubleshooting, patching, and performance analysis.
- Working knowledge of cloud and enterprise networking, including VPCs/VNets, subnets, routing, VPNs, load balancing, DNS, and firewalls.
- Experience implementing monitoring, logging, alerting, and observability for production platforms.
- Understanding of platform security, identity and access management, vulnerability remediation, and secure configuration practices.
- Ability to independently lead complex technical initiatives from design through production deployment.
- Strong technical documentation, communication, collaboration, and organizational skills.
- Bachelor's degree in computer science or a related field, or equivalent practical experience.
Preferred Qualifications
- Experience supporting commercial, government, regulated, classified, or air-gapped cloud environments.
- Experience with both Microsoft Azure and AWS, including Azure Government or AWS GovCloud.
- Experience with GitOps, Helm, policy as code, secrets management, service catalogs, or internal developer portals.
- Experience designing or supporting internal developer platforms and standardized paved-path tooling.
- Knowledge of private cloud and virtualization platforms such as VMware, Hyper-V, or KVM.
- Experience defining service-level indicators, service-level objectives, incident response practices, and reliability improvements.
- Experience supporting hybrid-cloud architectures spanning public cloud, private cloud, and on-premises infrastructure.
- Experience within aerospace, defense, manufacturing, or other regulated environments.
- Relevant cloud, Kubernetes, Terraform, or platform-engineering certifications.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://jobs.lever.co/shieldai/38e5f38c-8d3f-4210-8310-8688182dc182