# Senior Vision Language Model Engineer

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Vision-Language-Model-Engineer_JR2020818?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_d483ea59-31f

## Description

NVIDIA is seeking a senior vision language model engineer to design and build agentic data and training workflows for Autonomous Vehicles, Robotics, and Medical applications. The successful candidate will have the opportunity to work on cutting-edge projects and contribute to the development of NVIDIA's dataset search platforms for physical AI developers.

The role involves:

- Partnering with researchers to develop and evaluate prototypes of latest models, such as VLMs and VLAs, for video search, video understanding, and more

- Designing and implementing agentic data workflows that automate data discovery, labeling, evaluation, and retraining to maximize development velocity

- Building, curating, and maintaining high-quality multimodal datasets tailored for end-to-end physical AI problems

- Exploring and productizing new data sources, including simulation and synthetic data

- Using agentic AI workflows across the full applied research lifecycle

- Collaborating with research, model development, performance, and product teams

- Contributing to NVIDIA Cosmos Dataset Search and other core NVIDIA platforms and products

The ideal candidate will have:

- A PhD with 4+ years, MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field

- A strong background in modern deep learning, including transformer-based architectures, video modeling, and multimodal VLM/VLA or foundation models

- Excellent experience training and deploying deep learning models on real-world datasets

- Excellent experience with Python and at least one deep learning framework

- Current knowledge of the latest research on image and video search in autonomous vehicles, healthcare, robotics, or related physical AI applications

- Fluent with agentic AI workflows across the full applied research lifecycle

- Clear and effective communication skills, with experience working well in a dynamic, product- and research-focused team

NVIDIA offers equity and benefits to its employees.

## Skills

### Required
- Computer Science
- Computer Engineering
- Deep Learning
- Python
- Transformer-based architectures
- Video Modeling
- Multimodal VLM/VLA or Foundation Models

### Nice to have
- Publishing in top-tier conferences
- Patents in video retrieval
- Strong coding architecture skills
- Experience in robotic systems

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Vision-Language-Model-Engineer_JR2020818?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
