Description
NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. The successful candidate will be a key player in designing and implementing large-scale compute runs, working with the latest Accelerated computing and Deep Learning software and hardware platforms.
Responsibilities:
- Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting
- Manage Linux job/workload schedules and orchestration tools
- Develop and maintain continuous integration and delivery pipelines
- Develop tooling to automate deployment and management of large-scale infrastructure environments
- Deploy monitoring solutions for servers, network, and storage
- Perform troubleshooting from bare metal to application level
- Develop and document standard methodologies to share with internal teams
- Support Research & Development activities and engage in POCs/POVs for future improvements
Requirements:
- A degree in Computer Science, Engineering, or a related field and 8+ years of experience
- Knowledge of HPC and AI solution technologies
- Experience with job scheduling workloads and orchestration tools such as Slurm, K8s
- Excellent knowledge of Windows and Linux networking and internals
- Experience with multiple storage solutions such as Lustre, GPFS, Weka.io
- Python programming and bash scripting experience
- Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/chef
- Deep knowledge of Networking Protocols like InfiniBand, Ethernet
- Deep understanding and experience with virtual systems
- Familiarity with cloud computing platforms
Preferred Qualifications:
- Knowledge of CPU and/or GPU architecture
- Knowledge of Kubernetes, container-related microservice technologies
- Experience with GPU-focused hardware/software (DGX, Cuda)
- Experience with RDMA (InfiniBand or RoCE) fabrics
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Switzerland-Remote/Senior-HPC-AI-Cluster-Engineer_JR2018385