Description
We are seeking a highly skilled and hard-working Senior Test Developer/test engineer to join our multifaceted Enterprise Software QA team.
This role offers an outstanding opportunity to contribute to the design, construction, optimization, and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings.
Responsibilities:
- Collaborate with development teams on test plans for all layers of the software stack for cloud infrastructure, including execution, reviews, failure analysis, and assessing overall quality and risk.
- Work with customer project managers on software issues, including technical feedback from OEMs and CSPs, and develop key benchmarks to track execution and deploy process improvements to enhance efficiency.
- Utilize AI skills to expedite test scope, test planning, execution, and automation workflows.
- Lead NVIDIA Cloud and Data Center bring-up activities, involving validation, reporting, working with engineering to debug issues, providing design input, and adding coverage in different areas.
- Design, develop, and maintain CI/CD pipelines for continuous testing in cloud environments as needed.
- Perform performance, scalability, and reliability testing of cloud services.
- Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud.
- Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability.
- Work with various partner teams to ensure availability of clusters for testing and take the lead in resolving all issues.
- Collaborate with teams to ensure quality of cloud products delivered, focusing on critical areas like security, storage, workloads, performance on latest software and firmware components.
Requirements:
- A Master's or Ph.D. in Computer Science or a related field, or equivalent experience.
- Experience with AI development tools used in creating test cases, automating test cases, code coverage, and triaging.
- 8+ years of hands-on experience in cluster management and related tools, including Docker Containers, Slurm, Kubernetes, and Ansible.
- 2+ years of strong experience with cloud infrastructure platforms like AWS, Azure, Google, OCI Cloud.
- Hands-on experience with network, storage, security, cluster configuration, and debugging, as well as cloud infrastructure management tools like Terraform and Ansible.
- Expertise in administering, operating, and configuring Kubernetes.
- Experience in CI/CD tools such as GitLab and Jenkins, and the GitOps model.
- Proficiency in various monitoring tools: Prometheus, Grafana, CloudWatch, and Thanos.
- Proficiency in debugging issues involving networks, DHCP, DNS, HTTP, Linux, and containers.
Preferred Qualifications:
- Familiarity with 'Base Command Manager' for managing and monitoring high-performance computing.
- Experience in writing automation for web applications using tools like Selenium, Playwright.
You will also be eligible for equity and benefits.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Software-Development-Engineer-in-Test_JR2020431