Description
Production engineering is a field that involves crafting, building, and maintaining large-scale production systems with high efficiency and availability.
As a Senior Storage Production Engineer at NVIDIA, you will play a critical role in ensuring that our internal and external-facing GPU cloud services meet reliability and uptime goals.
Responsibilities:
- Design, implement, and support large-scale storage clusters, ensuring scalability, high availability, and data integrity.
- Develop and maintain storage monitoring, logging, and alerting systems to ensure proactive detection and resolution of performance issues.
- Work with AI/ML workloads to improve storage architectures for low-latency access, efficient caching, and high-throughput performance.
- Improve the lifecycle of storage services – from inception and design to deployment, operation, and continuous optimization.
- Maintain production storage infrastructure by supervising availability, latency, and system health, leveraging predictive analytics and AI-driven automation.
- Optimize storage efficiency through compression, deduplication, tiering strategies, and intelligent workload placement.
- Scale storage systems sustainably using AI/ML-driven automation, policy-based tiering, and dynamic data migration techniques.
- Practice sustainable incident response and blameless root cause analysis.
Requirements:
- BS degree or equivalent experience in Computer Science, Storage Systems, or a related technical field with 8+ years of practical experience.
- Experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise-grade storage systems.
- Solid understanding of block, file, and object storage technologies, including their scalability, reliability, and performance characteristics and standard processes.
- Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
- Expertise in algorithms, data structures, complexity analysis, software design, and automating maintenance of large-scale Linux-based storage systems.
- Experience in one or more of the following: C/C++, Java, Python, Go, NodeJS, and Bash for storage automation, monitoring, and performance tuning.
- Hands-on experience with infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform for automating storage deployments.
Benefits:
- Equity
- Benefits package
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Storage-Production-Engineer---DGX-Cloud_JR2021335