# Deep Learning Engineer

**Company**: NVIDIA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Shanghai/Software-Engineering-Intern--DLFW-Comms---2027_JR2025701?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_dd10aaa9-499

## Description

NVIDIA is seeking a motivated Deep Learning engineer to integrate advanced communication technologies into AI stacks like PyTorch, vLLM, SGLang, TRT-LLM, and veRL.

You will work with the team that developed communication libraries, such as NCCL and NVSHMEM, for scaling Deep Learning applications. Your customers will have diverse multi-GPU needs, ranging from training on scales up to 100K GPUs to inference at microsecond latency.

**Responsibilities:**

- Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production.

- Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities.

- Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms.

- Conduct in-depth research to achieve SOL GPU performance.

- Build fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.

- Collaborate with a very dynamic team across multiple time zones.

**Requirements:**

- Pursuing a M.S. or Ph.D. in CE/CS/EE with a strong background in communication, kernel authoring, and/or AI training/inference.

- Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe).

- Solid understanding of LLM models and parallelisms.

- Adaptability and passion to learn new areas and tools.

- Flexibility to work and communicate effectively.

**Preferred Qualifications:**

- Development experience with frameworks such as PyTorch, JAX, TRT-LLM, vLLM, SGLang, or veRL.

- Experience with DL communication patterns such as Expert Parallelism (EP), TP, DP & PP.

- Experience with CUDA kernel optimization and profiling.

- Experience with large-scale training or production inference stack.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.

## Skills

### Required
- Python
- C++
- CUDA
- Deep Learning
- Communication Libraries

### Nice to have
- PyTorch
- JAX
- TRT-LLM
- vLLM
- SGLang
- veRL
- Expert Parallelism
- CUDA kernel optimization

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Shanghai/Software-Engineering-Intern--DLFW-Comms---2027_JR2025701?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
