# Senior Deep Learning Framework Communications Engineer

**Company**: NVIDIA
**Location**: Santa Clara
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Deep-Learning-Framework-Communications-Engineer_JR2011908?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_77550b51-62d

## Description

We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect -- for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community.

**Responsibilities:**

- Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production

- Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models.

- Improve AI compilers to hide communications or perform automatic fusion.

- Conduct in-depth AI workload performance characterization on multi-GPU clusters.

- Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.

- Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms.

- Influence the roadmap of communication libraries - NCCL & NVSHMEM.

- Collaborate with a very dynamic team across multiple time zones.

**Requirements:**

- B.S, M.S. or PHD in Computer Science, or related field (or equivalent experience) with 5+ software engineering and HPC/AI experience

- Development or integration experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang

- Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe)

- Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile)

- Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems)

- Understanding of HPC/AI communication concepts (1-sided v 2-sided communication, elasticity, resiliency, topology discovery, etc)

- Adaptability and passion to learn new areas and tools

- Flexibility to work and communicate effectively across different teams and timezones

**Nice to Have:**

- Experience with parallel programming on at least one communication runtime (NCCL, NVSHMEM, MPI). Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)

- Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc). Experience with programming for compute & communication overlap in distributed runtimes

- Experience with AI compiler pattern matching and lowering. Solid understanding of memory hierarchy, consistency model, and tensor layout

## Skills

### Required
- PyTorch
- TRT-LLM
- vLLM
- SGLang
- JAX
- NCCL
- NVSHMEM
- CUDA
- Python
- C++
- Triton
- cuTe
- torch.compile
- performance benchmarking
- AI clusters
- HPC/AI communication concepts
- 1-sided v 2-sided communication
- elasticity
- resiliency
- topology discovery

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Deep-Learning-Framework-Communications-Engineer_JR2011908?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
