# Senior HPC Performance Engineer

**Company**: NVIDIA
**Location**: Germany, UK, Poland, Switzerland
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Germany-Remote/Senior-HPC-Performance-Engineer_JR2001201?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_32d8ab39-566

## Description

## Senior HPC Performance Engineer

We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs.

### What you will be doing:

- Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.

- Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack

- Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available

- Triage and root-cause performance issues reported by our customers

- Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information

- Collaborate with a very dynamic team across multiple time zones

### What we need to see:

- M.S. (or equivalent experience) or PHD in Computer Science, or related field with relevant performance engineering and HPC experience

- 3+ yrs of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)

- Experience conducting performance benchmarking and triage on large scale HPC clusters

- Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)

- Implement micro-benchmarks in C/C++, read and modify the code base when required

- Ability to debug performance issues across the entire HW/SW stack. Proficient in a scripting language, preferably Python

- Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible, Docker)

- Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively across different teams and timezones

### Ways to stand out from the crowd:

- Practical experience with Infiniband/Ethernet networks in areas like RDMA, topologies, congestion control

- Experience debugging network issues in large scale deployments

- Familiarity with CUDA programming and/or GPUs

- Experience with Deep Learning Frameworks such PyTorch, TensorFlow

## Skills

### Required
- parallel programming
- communication runtime
- performance benchmarking
- operating systems principles
- scripting language
- containers
- cloud provisioning
- scheduling tools

### Nice to have
- Infiniband/Ethernet networks
- CUDA programming
- Deep Learning Frameworks

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Germany-Remote/Senior-HPC-Performance-Engineer_JR2001201?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
