# Product Evaluations Lead - Gen AI Software

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Work arrangement**: hybrid
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Product-Evaluations-Lead---Gen-AI-Software_JR2021025?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_af344496-dde

## Description

NVIDIA is seeking a highly analytical Product Evaluations Lead to own the evaluation, measurement, and go-forward analysis of our Generative AI Software, including the Nemotron family, as they are adopted by our most strategic enterprise ISV partners.

The successful candidate will drive end-to-end evaluation and analysis of NVIDIA's GenAI libraries, primarily Nemotron models, and translate complex model behavior into clear go/no-go recommendations and prioritized next steps.

Key responsibilities include:

- Developing and owning rigorous evaluation frameworks, partner experimentation, and benchmarks that measure model quality, reasoning, and agentic capabilities against partner requirements and real-world enterprise use cases.

- Sharing learnings and impact through regular readouts, dashboards, and progress tracking on evaluation results, adoption status, and recommended next steps to Product, Engineering, Research, and leadership stakeholders.

- Providing analytical judgment under high uncertainty, balancing model quality, risk, and impact from incomplete or noisy signals to inform high-stakes release and integration decisions.

- Partnering with research scientists and product teams to translate emerging model capabilities into measurable release criteria, and identifying the data and signals that feed the model-improvement flywheel.

- Representing partner and product evaluation needs to internal teams and contributing to the product roadmap by synthesizing cross-partner and cross-industry patterns captured from strategic engagements.

Requirements:

- 12+ years of experience in data science, model evaluation, experimentation, or analytics roles, with a focus on measuring AI/ML or GenAI systems.

- Master's or PhD in a quantitative field (e.g., Data Science, Statistics, Computer Science, Economics) or equivalent experience.

- Proven track record leading model evaluation and experimentation that directly informed high-stakes, go/no-go product or model-release decisions.

- Strong hands-on skills in Python and statistical/experimental methods, including A/B testing, causal measurement, and metric design and failure analysis.

- Direct experience evaluating and benchmarking LLMs or GenAI systems, including reasoning, agentic, and/or multimodal capabilities, and building high-quality evaluation datasets and human-evaluation strategies.

NVIDIA offers equity and benefits, including access to comprehensive benefits programs.

## Skills

### Required
- Python
- statistical/experimental methods
- A/B testing
- causal measurement
- metric design
- failure analysis
- LLMs
- GenAI systems
- evaluation frameworks
- benchmarks

### Nice to have
- reward models
- data flywheels
- evaluation signals
- LLM capability evaluation
- benchmarking

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Product-Evaluations-Lead---Gen-AI-Software_JR2021025?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
