Description
We are seeking a hands-on technical leader to build and lead a high-performance engineering organisation that architects, delivers, and operates production-grade software systems at global scale.
You will be responsible for transforming enterprise IT operations from manual, reactive workflows into fully automated, AI-driven platforms that scale with NVIDIA’s hyper-growth.
Responsibilities:
- Architect and ship agentic AI systems using LLM-based agents, tool calling, RAG, and orchestration frameworks delivering production-grade AI-assisted operations across enterprise IT domains including employee support, endpoint services, and IT support operations.
- Design and deploy autonomous AI agents that execute complex, multi-step enterprise workflows end-to-end coordinating approvals, vendor handoffs, cross-system data reconciliation, and exception handling with human-in-the-loop controls delivering measurable improvements in availability, cycle time, cost, and compliance.
- Engineer robust integration and automation platforms spanning ServiceNow, ERP and procurement systems, endpoint-management platforms, Own the full stack infrastructure, data pipelines, APIs, and user-facing applications.
- Set the engineering standard through hands-on technical leadership co-authoring production code, conducting rigorous code reviews, and personally driving system design for the most critical components.
- Recruit, develop, and retain top-tier engineering talent. Build a high-performing team culture grounded in engineering excellence, ownership, and continuous delivery.
- Define and execute a multi-quarter technical roadmap for automation and agentic operations across enterprise IT, with each initiative tied to quantifiable business outcomes (cost reduction, throughput, SLA improvement, headcount avoidance).
- Drive disciplined execution,project prioritization, milestone tracking, capacity planning, and on-time delivery,while maintaining engineering velocity in a fast-moving environment.
- Own talent strategy for the team, including hiring pipelines, performance calibration, and career development that builds a deep bench of engineering leaders.
Requirements:
- Bachelor's or Master's degree in a related field, or equivalent experience
- 10+ overall years of hands-on software engineering experience, with deep expertise in at least one of: Infrastructure, SRE, DevOps, or Production Engineering. 5+ years leading engineering teams, with direct experience hiring, growing, and managing IT engineers.
- Demonstrated ability to build engineering teams from zero and scale them in a high-growth, high-ambiguity environment.
- Deep expertise in designing and shipping production software systems,including integrations, automation platforms, and data pipelines,for complex enterprise operations at scale.
- Track record of modernizing enterprise IT operations platforms (e.g., asset management, endpoint services, IT supply chain, infrastructure operations) and deploying agentic AI into production,including multi-step autonomous execution, human-in-the-loop safeguards, exception handling, and governance frameworks with measurable business outcomes.
- Production-grade proficiency with infrastructure-as-code, CI/CD, containerization (Kubernetes), and cloud platforms (AWS, GCP, or Azure).
- Experience with monitoring and observability tools (Prometheus, Grafana, Datadog, PagerDuty, or similar).
- Fluent in Python, Go, or equivalent languages,able to architect, write, and review production-quality code, not just scripts.
- Executive-level communication skills with the ability to influence technical direction across engineering, product, and senior leadership.
- Proven ability to translate complex technical capabilities into quantifiable business value and present to VP/C-level audiences.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Manager--Software-Engineering---Agentic-IT-Operations_JR2023714