Description
CoreWeave is hiring an IT Systems Engineer (Cloud Infrastructure) to build and operate the GCP, Azure, and AWS platforms that power our internal IT and enterprise systems.
The role covers cloud infrastructure, enterprise IT, and the operational reliability of the platforms those systems run on.
You will work hands-on across our internal cloud foundations: landing zones, hybrid connectivity to data centers, and the infrastructure behind identity, collaboration, IT operations, and internal tooling.
The work is execution-heavy. You will also mentor junior engineers and contribute to team-level technical decisions.
Responsibilities
Cloud foundations and architecture (GCP, Azure and AWS)
- Provision and manage cloud resources within established landing zones, including compute, storage, IAM, and workload configuration.
- Manage IAM, service accounts, Workload Identity Federation, and Secret Manager for internal workloads, aligned with security baselines established by the cloud council.
- Build and maintain Infrastructure as Code using Terraform, including GCP-specific modules for project factory and IAM.
- Partner with Security and Networking teams to implement approved zero-trust and defense-in-depth patterns for internal cloud services.
- Contribute to cloud migrations and integrations tied to company growth, including acquisitions and platform consolidations.
Infrastructure you'll own
- You will build and operate the cloud infrastructure behind CoreWeave's internal IT platforms and enterprise tooling.
- Core enterprise platforms in GCP projects, multi-tenant Azure, and AWS.
- Shared internal services: VDI, bastion hosts, internal APIs, job runners, and internal tooling.
- GCP-native workloads on GKE, Cloud Run, or Compute Engine.
Operations, reliability, and automation
- Maintain the reliability of internal cloud environments, including performance, capacity, cost, and security posture.
- Set up and maintain monitoring, alerting, and logging across GCP, Azure, and AWS environments.
- Implement, test, and maintain backup, disaster recovery, and failover configurations based on defined recovery requirements.
- Identify bottlenecks and propose fixes across performance, security, and scalability.
- Build automation for common workflows using Cloud Build, GitHub Actions, or similar pipelines.
- Contribute to CI/CD pipelines for infrastructure deployment and management.
- Troubleshoot production issues, participate in incident response, and contribute to root-cause analysis and corrective actions.
Mentorship and collaboration
- Support and mentor less-experienced engineers through troubleshooting, documentation, code reviews, and knowledge sharing.
- Be a reliable technical resource for the team.
- Work with engineers across teams on system integrations.
- Explain technical ideas and trade-offs to engineers and non-engineers alike.
Requirements
- 7–10 years of experience as a Systems, Cloud, or Infrastructure Engineer supporting production environments.
- Strong hands-on GCP experience, including networking, IAM, compute, storage, security, and observability in multi-project environments.
- Working experience with at least one additional public cloud platform, preferably Azure or AWS, and the ability to apply equivalent infrastructure concepts across cloud providers.
- Strong hands-on experience using Terraform to provision and manage cloud infrastructure, including reusable modules, remote state, plan review, troubleshooting, and CI/CD-based deployment.
- Working knowledge of cloud networking, including VPCs/VNets, subnets, routing, DNS, firewalls, load balancers, peering, VPNs, and private connectivity.
- Experience managing IAM, service accounts, workload identities, secrets, and least-privilege access.
- Experience using Git-based development workflows, pull requests, code reviews, and automated validation.
- Experience operating Linux and Windows workloads in the cloud, including deployment, configuration, patching, hardening, and troubleshooting.
- Experience implementing monitoring, logging, alerting, dashboards, and operational runbooks for cloud infrastructure and workloads.
- Proficiency in at least one scripting language, such as Python, Bash, or PowerShell.
- Demonstrated ability to troubleshoot production issues across infrastructure, networking, IAM, operating systems, and application dependencies.
- Clear communication of technical decisions, risks, and trade-offs to technical and non-technical stakeholders.
Nice to Have
- Experience using AI-assisted development tools to accelerate infrastructure engineering, scripting, troubleshooting, and documentation.
- Familiarity with agentic AI concepts and experience building AI-enabled, end-to-end automation or operational workflows.
- Experience integrating AI models or agents with enterprise systems, APIs, cloud services, and automation platforms.
What We Offer
- Base salary range: $182,000 to $242,000.
- Comprehensive benefits program, including medical, dental, and vision insurance, company-paid Life Insurance, short and long-term disability insurance, flexible spending account, health savings account, tuition reimbursement, employee stock purchase program, mental wellness benefits, family-forming support, paid parental leave, flexible childcare support, 401(k) with a generous employer match, flexible PTO, catered lunch, and a casual work environment.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/coreweave/jobs/4709849006