New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Software Development Engineer in Test

NVIDIA
Apply →
senior full-time Santa Clara, CA

First indexed 2 Jul 2026

Description

We are seeking a highly skilled and hard-working Senior Test Developer/test engineer to join our multifaceted Enterprise Software QA team.

This role offers an outstanding opportunity to contribute to the design, construction, optimization, and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings.

Responsibilities:

  • Collaborate with development teams on test plans for all layers of the software stack for cloud infrastructure, including execution, reviews, failure analysis, and assessing overall quality and risk.
  • Work with customer project managers on software issues, including technical feedback from OEMs and CSPs, and develop key benchmarks to track execution and deploy process improvements to enhance efficiency.
  • Utilize AI skills to expedite test scope, test planning, execution, and automation workflows.
  • Lead NVIDIA Cloud and Data Center bring-up activities, involving validation, reporting, working with engineering to debug issues, providing design input, and adding coverage in different areas.
  • Design, develop, and maintain CI/CD pipelines for continuous testing in cloud environments as needed.
  • Perform performance, scalability, and reliability testing of cloud services.
  • Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud.
  • Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability.
  • Work with various partner teams to ensure availability of clusters for testing and take the lead in resolving all issues.
  • Collaborate with teams to ensure quality of cloud products delivered, focusing on critical areas like security, storage, workloads, performance on latest software and firmware components.

Requirements:

  • A Master's or Ph.D. in Computer Science or a related field, or equivalent experience.
  • Experience with AI development tools used in creating test cases, automating test cases, code coverage, and triaging.
  • 8+ years of hands-on experience in cluster management and related tools, including Docker Containers, Slurm, Kubernetes, and Ansible.
  • 2+ years of strong experience with cloud infrastructure platforms like AWS, Azure, Google, OCI Cloud.
  • Hands-on experience with network, storage, security, cluster configuration, and debugging, as well as cloud infrastructure management tools like Terraform and Ansible.
  • Expertise in administering, operating, and configuring Kubernetes.
  • Experience in CI/CD tools such as GitLab and Jenkins, and the GitOps model.
  • Proficiency in various monitoring tools: Prometheus, Grafana, CloudWatch, and Thanos.
  • Proficiency in debugging issues involving networks, DHCP, DNS, HTTP, Linux, and containers.

Preferred Qualifications:

  • Familiarity with 'Base Command Manager' for managing and monitoring high-performance computing.
  • Experience in writing automation for web applications using tools like Selenium, Playwright.

You will also be eligible for equity and benefits.