# Senior Software Engineer

**Company**: NVIDIA
**Location**: Bengaluru
**Experience**: senior
**Job type**: full-time
**Category**: Engineering
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Software-Engineer_JR2024294?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_09304a27-af2

## Description

We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform.

This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations.

**Responsibilities:**

- Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale.

- Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams.

- Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals.

- Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation.

- Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams.

- Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness.

- Provide technical leadership and mentorship, influence engineering standards and architecture decisions, and deliver measurable improvements in reliability, MTTR, operational toil, engineering productivity, and infrastructure efficiency.

**Requirements:**

- Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience, with 10+ years of software engineering, SRE, infrastructure, or distributed-systems experience and demonstrated technical leadership.

- Strong software engineering expertise in Go, Python, or equivalent languages, with experience designing and building production-grade distributed systems, platform services, APIs, and automation.

- Proven experience owning complex software/platform initiatives across multiple teams or infrastructure domains, from architecture and implementation through adoption and measurable impact.

- Deep understanding of distributed systems, event-driven architectures, microservices, APIs, and high-throughput data processing, including technologies such as Kafka, NATS, gRPC, or equivalent.

- Strong experience with modern observability and telemetry platforms, including OpenTelemetry, Prometheus, VictoriaMetrics, Vector, Loki, Grafana, ClickHouse, or equivalent technologies.

- Strong SRE and infrastructure knowledge across Kubernetes/OpenShift, VMware, bare-metal, storage, networking, and/or cloud environments, with experience using Terraform, Ansible, or equivalent automation technologies.

- Demonstrated ability to solve ambiguous problems, influence technical direction without direct authority, mentor engineers, establish engineering standards, and deliver measurable operational and business outcomes.

**Nice to Have:**

- Experience building software and reliability platforms for large-scale on-premises infrastructure, particularly Storage, Compute, Networking, VMware, and Kubernetes/OpenShift.

- Deep understanding of storage and infrastructure telemetry, including IOPS, latency, NVMe health, SAN/NAS topology, block/object storage, and infrastructure failure domains.

- Experience building self-healing systems, automated remediation, predictive operations, or autonomous SRE capabilities.

- Production experience applying Generative AI, AIOps, LLMs, or Agentic AI to incident triage, RCA, operational intelligence, debugging, or remediation; experience with LangChain, LlamaIndex, AutoGen, or equivalent is a plus.

- Experience building high-performance platform services using FastAPI, gRPC, or equivalent technologies.

## Skills

### Required
- Go
- Python
- distributed systems
- event-driven architectures
- microservices
- APIs
- high-throughput data processing
- Kafka
- NATS
- gRPC
- OpenTelemetry
- Prometheus
- VictoriaMetrics
- Vector
- Loki
- Grafana
- ClickHouse
- Kubernetes
- OpenShift
- VMware
- bare-metal
- storage
- networking
- cloud environments
- Terraform
- Ansible

### Nice to have
- Generative AI
- AIOps
- LLMs
- Agentic AI
- LangChain
- LlamaIndex
- AutoGen
- FastAPI

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-Software-Engineer_JR2024294?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
