Description
NVIDIA is seeking an NCX Senior Engineer to join our AI Accelerator team, collaborating closely with strategic customers to implement and enhance groundbreaking AI workloads. You will deliver hands-on technical assistance for advanced AI deployments, intricate distributed systems, and ensure customers realize efficient performance from NVIDIA's AI platform across varied environments.
Responsibilities:
In this role, you will develop innovative solutions that advance AI infrastructure capabilities. You will directly influence customer success with breakthrough AI initiatives.
- Build and deploy custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed on NVIDIA reference architectures.
- Act as the main technical contact for strategic NCPs, offer remote and on-site support, troubleshoot complex production problems, and guide partner engineering teams on NVIDIA platform guidelines.
- Deploy and manage AI workloads across DGX Cloud, NCP data centers, and major CSP environments using Kubernetes, containers, and GPU scheduling systems aligned to NCP builds.
- Profile and tune large-scale training and inference workloads on NCP platforms. Implement observability and SLO/SLA monitoring. Lead detailed efforts to reduce latency, cost, and operational risk.
- Implement and expand NVIDIA reference architectures on partner platforms, develop integrations with partner control planes and customer environments, and ensure smooth API, data pipeline, and enterprise software connectivity.
- Build detailed implementation guides, runbooks, and post-mortem documentation that codify standard methodologies for running NVIDIA AI workloads at scale on NCP platforms.
Requirements:
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
- 8+ years of experience in customer-facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering, ideally supporting large-scale cloud or service provider environments.
- Strong expertise in Linux systems, distributed computing, Kubernetes, containers, and GPU scheduling on multi-tenant or service-provider platforms.
- Demonstrated AI/ML experience supporting large-scale training and inference workloads (e.g., LLMs, generative models, recommendation systems) in production or critically important environments.
- Solid programming skills in Python/Go, with hands-on experience using frameworks such as PyTorch or TensorFlow for training and serving.
- Demonstrated capability to collaborate with customer and partner engineering teams in fast-paced environments, guide intricate technical investigations, and bring issues to root cause and resolution.
- Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade-offs, and recommendations to both engineering and leadership audiences.
Nice to Have:
- Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, Triton, NIM, and NVIDIA networking technologies such as InfiniBand and RoCE.
- Direct experience collaborating with NVIDIA Cloud Partners, hyperscale CSPs, or managed AI cloud platforms, including implementation of NVIDIA reference architectures for AI infrastructure.
- Deep familiarity with MLOps and cloud-native practices: containerization, CI/CD pipelines, observability stacks (Prometheus, Grafana, OpenTelemetry), and GitOps workflows.
- Background in infrastructure as code (Terraform, Ansible, or similar) for repeatable deployment and configuration of GPU-accelerated clusters and NCP building blocks.
NVIDIA offers competitive salaries and a generous benefits package.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/NCX-Senior-Engineer_JR2021935-1