# Principal Software Engineer – Infrastructure

**Company**: NVIDIA
**Location**: Santa Clara, CA
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer---Infrastructure_JR2024352-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_cd91232c-471

## Description

We are looking for a Principal Software Engineer to join our Configuration Management team and define the future of enterprise infrastructure automation, configuration management, and operational platforms.

This is a deeply technical, hands-on leadership role; you will architect and build foundational software systems that manage infrastructure consistently across compute, storage, networking, data centers, cloud, and hybrid environments.

**Key Responsibilities:**

- Define the multi-year technical vision and architecture for IT infrastructure automation, configuration management, orchestration, and self-service platforms.

- Set the technical direction for infrastructure automation and configuration management across technologies such as Ansible Automation Platform, AWX, Salt or equivalent platforms.

- Establish the strategy for applying AI to configuration management, including configuration intelligence, drift and compliance analysis, change-risk identification, root-cause assistance, intelligent recommendations, and guarded automated remediation.

- Remain deeply hands-on by developing prototypes and production software, reviewing critical code and designs, and resolving the most challenging technical and scalability problems.

- Build secure and scalable integrations across infrastructure platforms, cloud services, CMDB, secrets management, observability, and enterprise data systems.

- Set and drive configuration management strategy across networking, storage, and compute domains.

- Ensure platforms meet enterprise requirements for availability, scalability, performance, security, disaster recovery, observability, and operational support.

- Act as a force multiplier by mentoring senior and staff engineers, raising engineering standards, facilitating architectural decisions, and developing technical leaders across teams.

- Partner with engineering and executive leadership to translate critical business challenges into technical strategy, prioritized roadmaps, and measurable business outcomes.

**Requirements:**

- 15+ years of progressive software engineering experience, with a sustained record of delivering complex, business-critical platforms and distributed systems.

- Bachelor’s or Master’s degree in Computer Science, Engineering, or a similar domain, or a Master’s degree or equivalent experience.

- Deep knowledge of enterprise configuration management or infrastructure automation using technologies such as Ansible Automation Platform, AWX, Salt or equivalent platforms.

- Deep hands-on expertise in Go, Python, Java, or a comparable systems programming language, including APIs, concurrency, distributed systems, testing, debugging, and performance engineering.

- Strong hands-on experience in Linux, Kubernetes, containers, cloud and hybrid infrastructure, CI/CD and Infrastructure as Code.

- Strong experience developing and implementing enterprise data and automation pipelines using databricks or comparable large-scale data platforms.

- Proven experience defining, owning, and evolving the architecture of large-scale infrastructure platforms operating across multiple teams, data centers, or cloud environments.

- Demonstrated ability to identify organization-wide business and technical challenges, establish a clear strategy, and drive implementation across teams.

- Outstanding communication and technical leadership skills, with a proven ability to build consensus, influence senior leaders, mentor experienced engineers, and lead through ambiguity.

**Nice to Have:**

- Experience in architecting, deploying and managing Ansible Automation Platform at enterprise scale.

- Experience developing platforms that manage large global infrastructure fleets across data centers, public clouds, compute, storage, and networking environments.

- A history of leading the development and enterprise-wide adoption of an AI-enabled configuration management, infrastructure automation, or autonomous operations platform.

- Experience using AI to solve infrastructure challenges such as configuration drift, compliance, change-risk analysis, incident diagnosis, capacity management, predictive operations, or automated remediation.

NVIDIA offers highly competitive salaries and a comprehensive benefits package.

## Skills

### Required
- Ansible Automation Platform
- AWX
- Salt
- Go
- Python
- Java
- Linux
- Kubernetes
- containers
- cloud and hybrid infrastructure
- CI/CD
- Infrastructure as Code
- databricks

### Nice to have
- AI-enabled configuration management
- infrastructure automation
- autonomous operations platform

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer---Infrastructure_JR2024352-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
