Description
NVIDIA is looking for an experienced HPC DevOps Engineer to help build supercomputers and HPC clusters of the future.
As a Senior HPC DevOps Engineer, you'll be a key player in groundbreaking advancements in artificial intelligence and GPU computing.
Responsibilities:
- Design, implement, and maintain large-scale HPC/AI clusters with state-of-the-art monitoring, logging, and alerting systems.
- Utilize and develop tools to manage infrastructure as code, ensuring scalable and repeatable deployments.
- Develop and maintain continuous integration and continuous delivery (CI/CD) pipelines to automate and streamline deployment processes.
- Develop automation scripts and tools to automate deployment, configuration management, and operational monitoring.
- Develop complex networking automations.
- Perform comprehensive troubleshooting from bare metal to application level, ensuring system reliability and efficiency.
- Serve as a technical resource, developing and sharing best practices with internal teams.
- Support R&D activities and engage in proof of concepts (POCs) and proof of values (POVs) for future improvements.
Requirements:
- B.Sc. in Computer Science, Engineering, or a related field
- 5+ years of experience
- Advanced proficiency in programming and scripting languages, with a solid understanding of object-oriented programming principles.
- Familiarity with Jenkins, Ansible, Puppet/Chef.
- Deep understanding of Kubernetes and container-related microservice technologies.
- Hands-on experience with event streaming or message queue technologies e.g. Apache Kafka.
- Experience with multiple storage solutions like Lustre, GPFS, ZFS, and XFS.
- Expertise with virtual systems (VMware, Hyper-V, KVM, Citrix).
- Familiarity with cloud platforms (AWS, Azure, Google Cloud).
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Israel-Tel-Aviv/Senior-HPC-DevOps-Engineer--NCS_JR2022973