Description
NVIDIA is seeking a Cloud Site Reliability Engineering Architect to work in its Infrastructure, Planning and Process (IPP) Cloud Infrastructure Team. The team provides cloud services that support various groups within NVIDIA, delivering unified CI/CD solutions and cloud-based software development.
The successful candidate will serve as an SRE Architect in the GPU Private Cloud team, used by thousands of NVIDIA employees globally for interactive development, centralized CI/CD, and QA testing. Key responsibilities include:
- Evaluating, identifying, and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA
- Architecting, implementing, and supporting end-to-end CI/CD systems using open-source and NVIDIA proprietary software
- Onboarding internal development teams to Private cloud infrastructure
- Identifying performance bottlenecks and optimizing the speed and cost efficiency of AI development and testing systems
- Leading software development projects and technically directing a team of engineers
- Resolving issues within software systems and crafting critical metrics using various analytics methods and dashboards
The ideal candidate will have:
- A BS or MS in Electrical Engineering, Computer Science, or a relevant field (or equivalent experience)
- 15+ years of systems software development experience, including at least 1 year dedicated to developing/exploring AI
- Experience maintaining cloud infrastructure and highly available production environments
- Strong programming and software development skills in Java, Python, Shell-script, and experience with distributed systems and REST APIs
- Experience working with SQL/NoSQL database systems, Docker containers, and Virtual Machines
- Knowledge of cloud technologies like OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, and Kafka
Equity and benefits are offered.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Engineer--Cloud-Site-Reliability-Engineering_JR2022907-1