New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Applied Research Engineer, Accelerator Algorithms

NVIDIA
Apply →
senior full-time Bengaluru

First indexed 16 Jul 2026

Description

As the world's leading accelerated computing company, NVIDIA is seeking a Senior Applied Research Engineer, Accelerator Algorithms to drive applied research and architecture for algorithms on NVIDIA's Programmable Vision Accelerator (PVA).

In this role, you will analyze real-world AV and physical AI workloads, identify workloads well suited for the PVA, and translate them into optimized algorithms.

Responsibilities:

  • Research and characterize real-world AV and physical AI workloads to identify algorithms that map well to PVA.
  • Develop and optimize PVA algorithms using instruction-level parallelism, memory-aware scheduling, efficient data movement, and vectorized execution.
  • Translate workload and algorithm insights into requirements for PVA hardware architecture, compilers, SDKs, profiling tools, and systems software.
  • Build prototypes, benchmarks, and performance models to evaluate PVA algorithm performance across current and future hardware architectures.
  • Leverage modern AI-assisted development tools and agentic workflows to accelerate workload analysis, algorithm creation, benchmarking, and performance exploration.
  • Partner with internal teams and customers to understand their systems and production workloads, and drive integration of PVA-accelerated algorithms into final products.
  • Publish, present, and communicate technical findings across research and engineering teams.

Requirements:

  • BS/MS or PhD in Computer Science, Electrical Engineering, Computer Engineering, Robotics, or a related field, or equivalent experience.
  • 12+ years of experience in applied research, programmable accelerator algorithm design, computer architecture, or high-performance computing.
  • Experience with DSP, SIMD, VLIW, fixed-point arithmetic, memory hierarchy, and low-level performance optimization.
  • Experience with HW/SW co-design, workload characterization, performance modeling, benchmarking, and bottleneck analysis.
  • Familiarity with emerging physical AI workloads such as VLA models, multimodal perception, sensor processing, autonomous systems, or robotics systems.
  • Strong programming skills in C++, Python, CUDA, or similar environments.
  • Excellent communication skills and the ability to explain complex technical tradeoffs clearly.
  • Strong ownership, technical judgment, and ability to operate in ambiguous problem spaces.

Nice to Have:

  • Hands-on experience with CUDA, ROS/ROS2, AV or robotics middleware, sensor processing frameworks, profiling tools, and heterogeneous compute pipelines.
  • Strong research track record in computer architecture, autonomous systems, or sensor processing.
  • Familiarity with safety-conscious development processes such as ISO 26262, IEC 61508, or similar standards.