Description
NVIDIA is seeking a Staff DB SRE to build the runtime foundation for NVIDIA's enterprise AI platforms, focusing on database infrastructure at scale. You will develop software systems, automation frameworks, and high-performance database services that power NVIDIA's AI workloads.
Responsibilities:
- Design and operate highly available database clusters with automated replication, failover, and disaster-recovery strategies.
- Drive database performance engineering, including query optimization, indexing strategies, and storage-engine tuning.
- Build self-service database lifecycle automation, including one-click cluster provisioning and zero-downtime upgrades.
- Bridge relational and AI-native data infrastructure, extending traditional database expertise into vector search and GPU-accelerated query engines.
- Compose and build software platforms that transform legacy database systems into modern and scalable architectures.
- Run vector and graph database services and query engines to handle AI/ML data workloads with ultra-low latency.
- Build automation frameworks for provisioning, schema evolution, scaling, and failover integrated directly into CI/CD workflows.
- Build developer-focused tooling for monitoring, profiling, and debugging database performance in real time.
- Contribute to architecture, coding standards, and guidelines for long-term platform evolution.
- Participate in on-call rotations to ensure flawless operation of critical database services.
Requirements:
- BS, MS, or PhD in Computer Science, Engineering, or a related field,or equivalent experience.
- 8+ years of Database engineering experience with deep expertise in database systems or distributed data platforms.
- Deep hands-on expertise with one or more major relational database engines, including replication topologies and high-availability architecture.
- Proven background in query optimization, data partitioning, and large-scale performance tuning.
- Experience building or operating managed database services or internal Database-as-a-Service platforms.
- Strong programming skills in Python, Go, with a track record of building production-grade systems.
- Demonstrable experience crafting high-performance, high-availability relational database services.
- Experience with container orchestration (Kubernetes) and cloud-native database deployment patterns.
- Hands-on experience establishing DevOps guidelines, e.g., CI/CD, monitoring, alerting, SLAs, capacity forecasting.
Nice to Have:
- Strong Kubernetes/Infrastructure as code and coding experience.
- Expertise in hybrid/multi-region database replication strategies for low-latency AI workloads.
- Demonstrable understanding of observability and performance profiling tools for complex data systems.
- Experience building an internal Database-as-a-Service offering.
- Hands-on experience with database migration tooling, schema evolution pipelines, and zero-downtime upgrade strategies.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Israel-Yokneam/Senior-Site-Reliability-Engineer_JR2023915