Description
NVIDIA is seeking a Machine Learning Engineer to help build the core of an autonomous, agentic platform that optimizes machine-learning models end-to-end. The platform optimizes the model itself (architecture, hyperparameters) and its implementation (the CUDA/Triton code it compiles to) across domains.
Responsibilities:
- Develop and advance a self-governing, agentic platform that optimizes AI models end-to-end , architecture, hyperparameters, and the GPU code they compile to.
- Leverage AI-native and agentic workflows to accelerate research, experimentation, evaluation, and deployment of AI systems.
- Establish and drive benchmarking frameworks that measure accuracy, latency, memory footprint, throughput, and cost , including head-to-head comparisons that prove the agent beats existing automated search.
- Design and deploy with strong consideration for reproducibility, AI safety, sandboxing, and compute-cost governance.
- Lead technical initiatives, mentor engineers, and foster a One Team culture through close collaboration across research, engineering, and product teams.
Requirements:
- Master's degree in Computer Science, AI, Electrical Engineering, or equivalent experience.
- 3+ years of experience building and deploying ML, LLM, or model-optimization systems.
- Strong Python skills and hands-on experience with PyTorch (or TensorFlow).
- Hands-on experience with automated experimentation , hyperparameter optimization, AutoML, or NAS.
- Experience building LLM-agent systems (reasoning, tool use, multi-step orchestration) and/or production ML pipelines and MLOps infrastructure.
- Proven technical leadership and mentoring experience, and strong problem-solving, communication, and teamwork skills.
Preferred qualifications:
- Hands-on experience with NVIDIA AI technologies such as NeMo, TAO, Triton, CUDA, NIM, and Nemotron.
- Experience building agentic AI systems with reasoning, tool use, and code generation.
- Expertise in optimization: evolutionary and quality-diversity search (e.g. MAP-Elites), Bayesian optimization, and multi-fidelity methods (Hyperband/ASHA).
- GPU performance work , CUDA/Triton kernels, torch.compile, operator fusion, quantization , and interest in inference-efficiency domains such as AI-RAN.
- Experience benchmarking AI systems for accuracy, latency, memory, reliability, and cost. A research track record (publications or credible reproductions) in AutoML, NAS, LLM agents, or optimization.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Vietnam-Ho-Chi-Minh-City/LLM-Engineer--Agentic-Researcher-Platform_JR2021274