# Senior LLM Agents Architect

**Company**: NVIDIA
**Location**: Yokneam
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Salary**: Competitive salaries and a comprehensive benefits package
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Israel-Yokneam/Senior-LLM-Agents-Architect_JR2018397?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_5bfe9de0-5a4

## Description

We are looking for a senior LLM Agents Architect to work hands-on with hardware architects, verification engineers, GPU performance experts, and software developers to build end-to-end agent flows that drive significant improvements in kernel optimization, architectural exploration, and developer efficiency.

Design and build agentic AI systems that generate, analyze, and optimize GPU compute kernels , targeting speed-of-light performance on NVIDIA hardware.

Collaborate with GPU architects and performance engineers to encode domain expertise , memory hierarchy trade-offs, occupancy tuning, instruction-level reasoning , into agent workflows that rival hand-tuned optimization.

Build automated performance forensics agents capable of ingesting large-scale simulation traces and Nsight profiler data to identify bottlenecks and propose architectural or software mitigations.

Partner with HW architects to develop agentic flows for GPU architectural studies , enabling rapid what-if analysis across micro-architecture configurations such as cache sizing, memory controller design, and compute unit scaling.

Explore agentic approaches to HW/SW co-design challenges, including replacing or augmenting graph-compiler functionality (e.g., TorchInductor) with LLM-driven optimization and code-generation pipelines.

Rapidly prototype and thoughtfully productize; integrate with internal services, utilize GPU capabilities, remove bottlenecks, and deliver fitting solutions.

Set up evaluation backbone using offline golden sets and online telemetry for confident iterations, cost control, and safe improvements.

Mentor and improve teams through insights in agent orchestration, prompting, RAG, observability, crafting documentation and playbooks for NVIDIA's teams.

## Skills

### Required
- applied ML/AI
- large-scale systems
- agentic or LLM-powered applications
- CUDA programming
- GPU architecture
- Nsight Compute
- Nsight Systems
- Python
- systems language (C++)
- tool use
- RAG pipelines
- model adaptation techniques

### Nice to have
- PyTorch compilation and lowering stack
- GPU graph compilers
- kernel fusion strategies
- auto-tuning frameworks
- performance engineering for HPC or GPU-accelerated workloads
- distributed processing
- multi-GPU workloads
- networking (e.g., NVLink, InfiniBand)
- frontier agentic coding tools (e.g., Claude Code, Codex, Cursor)

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Israel-Yokneam/Senior-LLM-Agents-Architect_JR2018397?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
