Description
We are seeking a Storage Production Engineer to join our team at NVIDIA. As a Storage Production Engineer, you will be responsible for designing, implementing, and supporting large-scale storage clusters, ensuring scalability, high availability, and data integrity.
Your primary focus will be on ensuring that our internal and external-facing GPU cloud services meet reliability and uptime goals. You will work closely with developers to make changes to the existing system while keeping an eye on capacity, latency, and performance.
Key responsibilities include:
- Designing, implementing, and supporting large-scale storage clusters
- Developing and maintaining storage monitoring, logging, and alerting systems
- Working with AI/ML workloads to improve storage architectures for low-latency access and high-throughput performance
- Improving the lifecycle of storage services, from inception to deployment and continuous optimization
- Maintaining production storage infrastructure and supervising availability, latency, and system health
- Optimizing storage efficiency through compression, deduplication, and tiering strategies
- Scaling storage systems sustainably using AI/ML-driven automation and policy-based tiering
- Practicing sustainable incident response and blameless root cause analysis
Requirements:
- BS degree or equivalent experience in Computer Science, Storage Systems, or a related technical field with 8+ years of practical experience
- Experience with distributed and high-performance storage solutions
- Solid understanding of block, file, and object storage technologies
- Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics
- Expertise in algorithms, data structures, and software design
- Experience with programming languages such as C/C++, Java, Python, Go, NodeJS, and Bash
- Hands-on experience with infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform
Nice to have:
- Deep understanding of extensive distributed storage systems and replication strategies
- Experience with Git, code review, pipelines, and CI/CD
- Strong debugging skills with a systematic problem-solving approach
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Australia-Remote/Senior-Storage-Production-Engineer---DGX-Cloud_JR2025191