New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
NVIDIA

Senior Software Engineer, Core Infrastructure Services - DGX Cloud

NVIDIA
Apply →
remote senior full-time

First indexed 5 Aug 2026

Description

NVIDIA is seeking an experienced software engineer to join the Cloud Foundations Automation team. The team builds and operates core infrastructure services that power NVIDIA's DGX Cloud and SuperPod deployments.

Job Overview: You will build and operate core infrastructure services that power NVIDIA's global AI infrastructure. You will architect and develop secure, scalable, and highly available cloud-native platform services. Your responsibilities will also include developing software for infrastructure orchestration, self-service workflows, and platform automation.

Key Responsibilities:

  • Build and operate core infrastructure services that power NVIDIA's global AI infrastructure.
  • Architect and develop secure, scalable, and highly available cloud-native platform services.
  • Develop software that enables infrastructure orchestration, self-service workflows, and platform automation.
  • Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management.
  • Build observability and security capabilities that improve the reliability and resilience of our infrastructure.
  • Partner with infrastructure and networking teams to deliver production services at scale.
  • Drive operational excellence through automation, monitoring, incident response, and continuous improvement.

Requirements:

  • BS or equivalent experience with 8+ years of relevant industry experience.
  • Strong proficiency in Python and Go, with experience building production-quality software.
  • Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST.
  • Experience with infrastructure automation (Terraform, Ansible), workflow orchestration (Temporal), and distributed systems using databases, Redis, and messaging platforms (Kafka, NATS, SQS).
  • Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms.
  • Strong Linux fundamentals with experience in observability (Prometheus, Grafana, OpenTelemetry, gNMI), networking (BGP, switching, routing, load balancing), and security (VPNs, firewalls, iptables/nftables).
  • Excellent problem-solving, communication, and collaboration skills.

Nice to Have:

  • Hands-on experience with network infrastructure including switches, routers, and firewalls.
  • Familiarity with InfiniBand, RDMA, and AI/HPC networking.
  • Experience with NetBox, Nautobot, or similar network source of truth platforms.
  • Contributions to open-source software.
  • Experience with public cloud platforms.

Benefits: You will also be eligible for equity and benefits.