Description
CoreWeave is The Essential Cloud for AI. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence.
As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers.
The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems.
This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.
Responsibilities
- Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.
- Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.
- Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.
- Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly.
- Translate incidents and hardware failure modes into software improvements that make the platform more resilient.
- Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale.
- Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship.
- Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden.
Requirements
- B.S., M.S., or PhD in Computer Science or related field, or equivalent experience.
- 8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases,particularly within large-scale datacenter and cloud environments.
- Expertise in Go and proven experience building REST/gRPC APIs for mission-critical platforms.
- Strong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
- Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams.
- Experience contributing to and collaborating with open source communities.
- Skilled in applying a data-driven approach to reliability, optimization, and continuous improvement.
- Excellent communicator able to work effectively with both technical and non-technical stakeholders.
- Hands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers.
- Track record of leading incident response, postmortems, and driving robust service reliability.
Nice To Have Skills
- Working knowledge of Kafka, ClickHouse and CRDB.
- DMTF, RedFish APIs, and GPU servers.
Compensation
The base salary range for this role is $207,000 to $275,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location.
Benefits
- Medical, dental, and vision insurance - 100% paid for by CoreWeave
- Company-paid Life Insurance
- Voluntary supplemental life insurance
- Short and long-term disability insurance
- Flexible Spending Account
- Health Savings Account
- Tuition Reimbursement
- Ability to Participate in Employee Stock Purchase Program (ESPP)
- Mental Wellness Benefits through Spring Health
- Family-Forming support provided by Carrot
- Paid Parental Leave
- Flexible, full-service childcare support with Kinside
- 401(k) with a generous employer match
- Flexible PTO
- Catered lunch each day in our office and data center locations
- A casual work environment
- A work culture focused on innovative disruption