Description
At NVIDIA, we are at the forefront of technological innovation, pushing the boundaries of AI and accelerated computing. Our team in Shanghai, China is looking for a Senior Site Reliability Engineer focused on Test Environment Management to join us. This is an opportunity to build and operate highly reliable test infrastructure, CI/CD systems, and environments that power validation of NVIDIA enterprise offerings.
What you’ll be doing
- Design, build, operate, and continuously improve reliable, scalable test environments and automation infrastructure that support validation of NVIDIA enterprise offerings.
- Own end-to-end CI/CD pipelines using GitLab CI, GitHub Actions, and ArgoCD (GitOps) , including pipeline design, reliability, performance, and progressive delivery of test workloads.
- Manage Software Bills of Materials (SBOMs): generation, continuous monitoring, vulnerability correlation, policy enforcement, and integration into CI/CD and release gate..
- Provision, scale, observe, and lifecycle-manage ephemeral and long-lived test environments (Kubernetes-based and hybrid) with strong emphasis on isolation, reproducibility, and rapid recovery.
- Define and drive reliability practices for test systems: SLIs/SLOs/error budgets for test environments and pipelines, toil reduction, chaos/resilience testing of infrastructure, and automated remediation.
- Collaborate closely with development and platform teams to triage environment and pipeline failures, perform root-cause analysis, verify fixes, and continuously harden test infrastructure.
- Apply AI/ML/Agentic techniques and internal tools to accelerate environment provisioning, flaky-test detection, capacity planning, anomaly detection, and overall Quality Assurance velocity.
What we need to see
- MS or PhD in Computer Science, related field, or equivalent experience with 8+ years of professional experience in Site Reliability Engineering, Test Environment Management, CI/CD platform engineering, or software testing infrastructure.
- Strong proficiency with Linux, shell scripting, and Python (or equivalent automation languages).
- Hands-on experience designing and operating CI/CD systems with GitLab CI and/or GitHub Actions, practical experience with ArgoCD (or equivalent GitOps tooling) for CD of applications and infrastructure. Solid background in containerization and orchestration (Docker, Kubernetes) and virtualization technologies.
- Deep understanding of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, toil elimination, and reliability engineering for complex distributed systems.
- Experience building and operating test environments (ephemeral, multi-tenant, or production-like) with focus on reliability, isolation, and rapid turnaround.
- Strong knowledge of QA principles and how test infrastructure enables high-quality software delivery.
- Comfort working with AI/LLM-related workloads and toolings. Excellent problem-solving skills, clear written and verbal communication, and the ability to collaborate across engineering teams.
- Self-motivated, proactive, and passionate about learning new technology at scale.
Ways to stand out from the crowd
- Experience operating large-scale Kubernetes platforms and GitOps workflows in production or high-stakes test environments.
- Background in software supply-chain security, SBOM tooling ecosystems, vulnerability management, and policy enforcement (OPA/Gatekeeper, Kyverno, etc.).
- Hands-on work with NVIDIA GPU hardware, multi-GPU environments, or accelerated computing infrastructure. Experience with parallel programming, high-performance computing, or large-scale AI model training/inference test harnesses.
- Track record of applying AI/observability techniques to detect flaky tests, optimize environment utilization, or automate root-cause analysis.
- Prior experience defining and driving reliability programs (error budgets, chaos engineering, capacity forecasting) for CI/CD or test platforms.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/China-Shanghai/Senior-Site-Reliability-Engineer-in-Test--SDET_JR2023203-1