# Distinguished Engineer, Storage – AI Cloud

**Company**: NVIDIA
**Location**: Santa Clara
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Salary**: Competitive salary and benefits package
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Distinguished-Engineer--Storage---AI-Cloud_JR2018037?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_efeba8ac-cde

## Description

As a Distinguished Engineer, Storage – AI Cloud at NVIDIA, you will lead the multi-year technical plan for AI Cloud Storage expansion across Neocloud Providers (NCPs). You will determine the reference architecture, capabilities, performance and durability SLOs, qualification methodology, and roadmap for high-performance file, object, and block storage that each NCP must offer to qualify for NVIDIA GPU allocation.

You will serve as the chief storage architect with deep hands-on involvement. Lead key reviews of storage builds and investigate root causes of complex production problems. Develop prototype reference implementations to minimize risks in new initiatives. Make final technical decisions on NCP storage deliveries using measurable SLOs. Apply AI tools heavily to amplify your technical influence throughout the program.

Define the standard for 'production-ready' in NCP storage, including durability and availability SLOs measured in 9s. Ensure sustained efficiency per TiB, observability, blast-radius containment, and reduced operational toil. Influence GPU delivery gating by requiring AI Cloud to accept GPU capacity only after verifying storage-focused ancillary services.

Develop and guide the architectural direction by working closely with collaborators in training, inference, and accelerated-computing product lines. Coordinate with site-reliability, operations, networking, and security colleagues. Work together with external cloud providers, neocloud operators, and storage vendors to align on a common architecture.

Develop the open-source path forward for AI storage. Establish and guide an open-source strategy that broadens the AI storage ecosystem. Advocate for a GitHub-first, security-first stance. Engage deeply with upstream open-source communities. Formalize the APIs, SDKs, and protocols allowing partners and the industry to build, integrate, and create with NVIDIA at the AI storage level.

Lead an engineering culture centered on AI tools. Regularly use modern AI coding and agentic tools in your daily tasks. Show what 10× engineering means at NVIDIA. Distribute patterns, prompts, and evaluation harnesses across the storage organization.

Partner with peer Distinguished and Principal storage architects across the organization to tackle the most difficult, long-term technical challenges. Make automation the only acceptable solution for infrastructure management tasks like live software upgrades, node and drive replacements, capacity rebalancing, cross-DC data movement, and dataset lifecycle. Establish root-cause analysis and corrective action rigor on every major incident. Design the storage layer for workloads spanning the next several GPU generations, including disaggregated inference with storage-backed KV caching, large-scale write-once-read-many inference patterns, exabyte regional object stores, and cross-DC dataset versioning and copy management.

Mentor and develop senior, principal, and distinguished engineers across the storage organization and nearby business units. Raise the technical bar broadly. Represent NVIDIA externally in standards bodies, open-source communities, customer briefings, and industry forums (FAST, SC, OCP, SNIA, Linux Storage Summit).

To succeed in this role, you will need:

- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field , or equivalent experience.

- A minimum of 18+ years of practical engineering experience in storage technology is needed. This involves extensive involvement with a high-performance parallel file system like Lustre, GPFS / Spectrum Scale, WEKA, VAST, BeeGFS, DAOS, or its equivalent, handling data at multi-petabyte scale. Candidates must also have wide-ranging expertise in object storage (S3 / Swift-class) and block storage (NVMe-oF, NVMesh-class, iSCSI).

- A track record of crafting and managing storage platforms at exabyte scale for performance-critical workloads , AI training, HPC, video, or hyperscale data lakes , including direct responsibility for durability, availability, and performance SLOs measured in 9s.

- Demonstrated ability to set technical strategy across business units and partner organizations. You have driven multi-year storage architectures adopted by multiple teams, vendors, or customers. You can point to measurable outcomes such as GPU utility lift, $/PB reduction, incidents eliminated, and time-to-bring-up compressed.

- You are 100% hands-on in engineering. You write and review production code yourself. When a bug requires it, you read Lustre, NFS, kernel, NVMe-oF, or SPDK source code. You also run scale tests or recovery drills personally instead of delegating.

- Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).

- Frequent daily use of advanced AI coding and autonomous tools, including specific examples showing how you accelerated building, coding, debugging, validation, and operations. Also, share your perspective on future trends.

- Excellent written and verbal communication. You can write a one-pager that aligns a VP. You can also write a six-pager that aligns an entire org. You can explain a deep technical trade-off to an SRE, a vendor CTO, and an internal customer in the same week.

- Comfort operating in a 24/7 production environment where storage incidents directly impact GPU revenue, with a security-first approach baked into every build.

To stand out from the crowd, consider highlighting:

- Proven background in designing or managing storage solutions for AI training or inference at 10k+ GPU scale, demonstrating clear improvements in GPU utilization or reducing I/O bottlenecks.

- Open-source contributions or maintainership in Lustre, NFS, SPDK, NVMe / NVMe-oF, CSI, Ceph, MinIO, RocksDB, or related projects.

- Built or led a disaggregated-inference or

## Skills

### Required
- storage technology
- high-performance parallel file systems
- object storage
- block storage
- NVMe-oF
- NVMesh-class
- iSCSI
- Lustre
- GPFS / Spectrum Scale
- WEKA
- VAST
- BeeGFS
- DAOS
- S3 / Swift-class
- Python
- C
- C++
- Rust
- Go
- Linux kernel storage and networking stacks
- RDMA / RoCE / InfiniBand
- NVMe
- page cache
- VFS
- multipath

### Nice to have
- AI coding and autonomous tools
- disaggregated-inference
- large-scale write-once-read-many inference patterns
- exabyte regional object stores
- cross-DC dataset versioning and copy management

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Distinguished-Engineer--Storage---AI-Cloud_JR2018037?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
