Description
We are seeking a CPU computing engineer to join our team in Shanghai. As a CPU computing engineer, you will craft and develop robust inferencing software that can be scaled to multiple platforms for functionality and performance.
Responsibilities:
- Craft and develop robust inferencing software that can be scaled to multiple platforms for functionality and performance
- Performance analysis, optimization, and tuning
- Closely follow academic developments in the field of artificial intelligence and feature update TensorRT and TensorRT Edge LLM
- Collaborate across the company to guide the direction of machine learning inferencing, working with software, research, and product teams
Requirements:
- Masters or higher degree in Computer Engineering, Computer Science, Applied Mathematics, or related computing-focused degree (or equivalent experience)
- 4+ years of relevant software development experience
- Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design
- Strong curiosity about artificial intelligence, awareness of the latest developments in deep learning like LLMs, generative models
- Experience working with deep learning frameworks like PyTorch
- Proactive and able to work without supervision
- Excellent written and oral communication skills in English
- Strong customer communication skills, powerfully motivated to provide highly responsive support as needed
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Beijing/Software-Engineer--LLM-Inference_JR2024698