# Senior Staff AI Platform Engineer

**Company**: NVIDIA
**Location**: Santa Clara
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/AI-Platform-Engineer_JR2014718-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_5fabe8ec-858

## Description

Apply for the Senior Staff AI Platform Engineer position at NVIDIA, where you will have the opportunity to build, support, and maintain the next generation of AI-powered enterprise products. As a key member of our team, you will collaborate with Cloud and AI/ML teams in a multifaceted and agile environment.

**Key Responsibilities:**

- Define and lead AI-native infrastructure roadmaps and cross-organizational initiatives.

- Architect and scale LLM/ML infrastructure across cloud-native clusters and on-premises hardware.

- Design and implement observability for infrastructure health and AI model performance.

- Build LLM-aware monitoring and leverage AI to improve incident response and reduce toil.

- Develop automation and tooling to ensure reliability, scalability, and developer self-services.

- Troubleshoot complex distributed systems, including deep Kubernetes and AI/ML scaling challenges.

- Drive AI-assisted engineering practices and mentor engineers to foster an AI-first culture.

- Partner with product engineering and internal business units to translate AI platform capabilities into reliable, scalable solutions that accelerate product development.

**Requirements:**

- 10+ years in cloud, platform, or SRE roles with relevant education or equivalent experience.

- Bachelors degree or equivalent experience.

- Strong Python and at least one systems language (C++, Go, or Rust), with proven distributed systems debugging expertise.

- Deep experience building and scaling distributed systems, including Kubernetes and bare-metal infrastructure.

- Strong observability design across infrastructure and AI workloads (metrics, logging, tracing, AI quality signals).

- Hands-on experience operating AI/ML platforms, including MLOps, model serving, and GPU-accelerated environments.

- Experience with infrastructure and application security practices, such as identity/auth, network segmentation, supply chain security, and vulnerability management in cloud-native environments.

- Practical use of AI-assisted development tools and coding agents in daily workflows.

- Solid foundation in data structures, algorithms, and complexity analysis.

- Excellent problem-solving, communication, and collaboration across multiple functions.

**Nice to Have:**

- Deep experience with AI/ML platforms (e.g., Hugging Face, Weights & Biases, NVIDIA NIM).

- Proven use of AI agents and LLM tooling to enhance observability, incident response, or developer productivity.

- Experience with artifact management, AI supply chain security, or trusted model distribution systems.

- Experience with AI-specific threat models (OWASP Top 10 for LLMs, model poisoning, adversarial inputs), experience with FedRAMP, SOC 2, or other compliance frameworks relevant to your environment, and red-teaming or security evaluation of LLM systems.

- Strong ownership demeanor with a structured, automation-first approach.

- Demonstrated impact driving AI-first engineering practices across teams.

## Skills

### Required
- Python
- C++
- Go
- Rust
- Kubernetes
- LLM/ML infrastructure
- Observability
- AI/ML platforms
- MLOps
- Model serving
- GPU-accelerated environments
- Infrastructure security
- Application security
- AI-assisted development tools
- Data structures
- Algorithms
- Complexity analysis

### Nice to have
- Hugging Face
- Weights & Biases
- NVIDIA NIM
- Artifact management
- AI supply chain security
- Trusted model distribution systems
- AI-specific threat models
- FedRAMP
- SOC 2
- Compliance frameworks
- Red-teaming
- Security evaluation

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/AI-Platform-Engineer_JR2014718-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
