Description

We are now looking for a Senior Deep Learning Software Engineer, LLM Performance! NVIDIA is seeking an experienced Deep Learning Engineer passionate about analyzing and improving the performance of LLM inference.

As a Senior Deep Learning Software Engineer, you will work with the deep learning community to implement the latest algorithms for public release in TensorRT LLM, VLLM, SGLang and LLM benchmarks. You will identify performance opportunities and optimize SoTA LLM models across the spectrum of NVIDIA accelerators, from datacenter GPUs to edge SoCs.

Your responsibilities will include:

Performance optimization, analysis, and tuning of LLM, VLM and GenAI models for DL inference, serving and deployment in NVIDIA/OSS LLM frameworks.

Scale performance of LLM models across different architectures and types of NVIDIA accelerators.

Scale performance for max throughput, minimum latency and throughput under latency constraints.

Contribute features and code to NVIDIA/OSS LLM frameworks, inference benchmarking frameworks, TensorRT, and Triton.

Work with cross-collaborative teams across generative AI, automotive, image understanding, and speech understanding to develop innovative solutions.

To succeed in this role, you will need:

Bachelors, Masters, PhD, or equivalent experience in relevant fields (Computer Engineering, Computer Science, EECS, AI).

At least 8 years of relevant software development experience.

Excellent Python/C/C++ programming, software design and software engineering skills.

Experience with a DL framework like PyTorch, JAX, TensorFlow.

Prior experience with a LLM framework or a DL compiler in inference, deployment, algorithms, or implementation will be a plus.

This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Deep-Learning-Software-Engineer--LLM-Performance_JR2016389