Description
NVIDIA's AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software vulnerabilities. We are looking for an Evaluation/ML-Systems Engineer to own how we measure the program.
You will play a critical role in separating real capability from anecdote. You will make our numbers mean something, and keep them meaning it as the program grows. Your benchmarks will define what progress means for the program. Every conclusion the program reaches should trace back to benchmarks, runs, and code you helped make reproducible.
Responsibilities:
- Evaluation infrastructure: Build the benchmarking and reproducibility systems we depend on.
- Metrics and protocols: Define the metrics and protocols we measure against.
- Traceability: Map every result to the code and runs that produced it.
- Evidence discipline: Keep findings reviewable and conclusions traceable.
Requirements:
- Bachelor's degree (or equivalent experience) with 5+ years in ML engineering or evaluation.
- Evaluation experience: Designing benchmarks, metrics, and statistically sound comparisons for ML systems.
- Measurement rigor: A careful, skeptical approach to metrics, baselines, and claims.
- Engineering skills: Solid Python engineering for shared infrastructure, including experiment tracking and data pipelines.
Nice to Have:
- Security evaluation: Exposure to evaluating security tooling or pipelines.
- Agentic systems: Experience measuring agent or LLM behavior.
- Community work: Contributions to public benchmarks or evaluation frameworks.
NVIDIA offers competitive salaries and a generous benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Evaluation-and-ML-Systems-Engineer--AI-Safety-and-Security-Engineering_JR2021888