Description
NVIDIA is looking for a Senior System Software Engineer with deep expertise in speech technologies to support enterprise and developer customers. This role involves hands-on technical engagement with customers to implement, troubleshoot, and optimize Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Audio Language Models (ALM), and Speech-to-Speech (S2S) systems in production environments.
Responsibilities:
- Work on cutting-edge GPU-accelerated AI systems deployed at scale
- Tackle challenging problems in real-time streaming audio processing and low-latency inference
- Troubleshoot and resolve complex issues across ASR, TTS, ALM, and S2S pipelines
- Model Integration: Work alongside Model researchers to transition ASR, TTS and S2S models from research to production readiness
- Develop Core Speech Services: Build and enhance C++ & python backend implementations for ASR, TTS, and S2S pipelines, leveraging CUDA for GPU acceleration
- Optimize Inference Performance: Improve streaming latency and throughput through advanced batching strategies, encoder caching, and multi-threaded pipeline optimizations
- Feature Development: Add new capabilities such as advanced voice activity detection, speaker diarization, decoder implementations (CTC, WFST, Flashlight), and text post-processing
- Client Libraries: Contribute to Python and C++ client SDKs and CLI tools for easy service integration
- Assist with API integration, SDK usage, model deployment, and performance optimization
- Provide advanced technical guidance to customers implementing speech technology solutions
Requirements:
- Masters or BE/BTech in Computer Science, computer architecture, or related field
- 6+ years of experience
- Excellent C++ & Python programming and software design skills, including debugging, performance analysis, and test design
- Experience with inference pipelines for LLM, Speech Recognition & Speech Synthesis
- Solid understanding of modern model architectures (Transformers, CNNs, RNNs)
- Excellent debugging abilities spanning multiple software (storage systems, kernels and containers)
- Experience building and deploying cloud services using HTTP REST, gRPC, Websockets and related technologies
- Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment
- Ability to work independently, define project goals and scope and manage your own development effort
- Knowledge of real-time streaming audio systems and low-latency architectures
- Experience with speech model fine-tuning or customization
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Pune/Senior-System-Software-Engineer--Speech-AI_JR2020087-1