Description
NVIDIA's accelerated computing platform is enabling significant advancements in large language models, but these models' scale and complexity create new challenges in computational efficiency. We're seeking a strong technical leader to drive a unified strategy for making LLMs more efficient from research through deployment.
You will lead a multidisciplinary effort to bring together model innovation, systems expertise, and hardware awareness. Your goal will be to ensure new capabilities are delivered within practical constraints of compute, memory, power, and cost.
Responsibilities:
- Lead cross-layer efforts to improve the efficiency of large language models across model architecture, training, and inference systems.
- Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure, identifying opportunities for model-system-hardware co-design.
- Establish a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment.
- Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence future model, software, and hardware roadmaps.
Requirements:
- MS or PhD degree, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
- 5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
- Strong understanding of LLM architectures, training and inference workloads, and tradeoffs between model quality, computational cost, memory footprint, latency, throughput, and power.
- Strong background in performance analysis, roofline modeling, workload characterization, benchmarking, and hardware-aware optimization.
- Proven ability to provide technical leadership and drive complex optimization projects from concept to production.
Preferred Qualifications:
- A track record of delivering measurable improvement in throughput, cost per token, energy per token, memory efficiency, or time to train.
- A first-principles approach to improving LLM efficiency: measure, model, optimize, and deliver.
- Familiarity with low-precision computation, quantization, sparsity, Mixture-of-Experts, long-context inference, and speculative decoding.
- Experience co-designing model architectures with training, inference, compiler, or hardware constraints.
- Experience influencing accelerator, system, or datacenter architecture based on future AI workload requirements.
You will also be eligible for equity and benefits.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-DL-Performance-Efficiency-Architect--_JR2024654