Description
NVIDIA is seeking a Principal Research Scientist to lead synthetic data generation across its frontier model efforts. You will define and build open-source libraries within the NVIDIA NeMo ecosystem to generate synthetic datasets for large language models like Nemotron.
Responsibilities:
- Build and scale data generation pipelines using LLM-based methods combined with automated quality evaluation.
- Pioneer data generation for agentic and tool-use training, including synthetic trajectories and multi-turn interactions.
- Advance multimodal synthetic data generation in partnership with NVIDIA's model teams.
- Develop and maintain open-source libraries and SDKs with clean APIs and strong documentation.
- Publish original research at top machine learning and AI conferences.
- Mentor scientists and engineers across the team.
Requirements:
- PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.
- 15+ years of engineering and research experience in synthetic data generation, generative modeling, or related areas.
- Deep technical understanding of LLMs, data pipelines, and inference frameworks.
- Proven track record of developing software libraries used by a broad developer community.
- Experience building scalable data pipelines for large-scale model training.
- Strong publication record at premier venues like NeurIPS, ICML, ICLR, or ACL.
Benefits:
- Equity
- Benefits package
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Research-Scientist--Synthetic-Data-Generation_JR2021918