Description
NVIDIA is seeking a highly analytical Product Evaluations Lead to own the evaluation, measurement, and go-forward analysis of our Generative AI Software, including the Nemotron family, as they are adopted by our most strategic enterprise ISV partners.
The successful candidate will drive end-to-end evaluation and analysis of NVIDIA's GenAI libraries, primarily Nemotron models, and translate complex model behavior into clear go/no-go recommendations and prioritized next steps.
Key responsibilities include:
- Developing and owning rigorous evaluation frameworks, partner experimentation, and benchmarks that measure model quality, reasoning, and agentic capabilities against partner requirements and real-world enterprise use cases.
- Sharing learnings and impact through regular readouts, dashboards, and progress tracking on evaluation results, adoption status, and recommended next steps to Product, Engineering, Research, and leadership stakeholders.
- Providing analytical judgment under high uncertainty, balancing model quality, risk, and impact from incomplete or noisy signals to inform high-stakes release and integration decisions.
- Partnering with research scientists and product teams to translate emerging model capabilities into measurable release criteria, and identifying the data and signals that feed the model-improvement flywheel.
- Representing partner and product evaluation needs to internal teams and contributing to the product roadmap by synthesizing cross-partner and cross-industry patterns captured from strategic engagements.
Requirements:
- 12+ years of experience in data science, model evaluation, experimentation, or analytics roles, with a focus on measuring AI/ML or GenAI systems.
- Master's or PhD in a quantitative field (e.g., Data Science, Statistics, Computer Science, Economics) or equivalent experience.
- Proven track record leading model evaluation and experimentation that directly informed high-stakes, go/no-go product or model-release decisions.
- Strong hands-on skills in Python and statistical/experimental methods, including A/B testing, causal measurement, and metric design and failure analysis.
- Direct experience evaluating and benchmarking LLMs or GenAI systems, including reasoning, agentic, and/or multimodal capabilities, and building high-quality evaluation datasets and human-evaluation strategies.
NVIDIA offers equity and benefits, including access to comprehensive benefits programs.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Product-Evaluations-Lead---Gen-AI-Software_JR2021025