Description
We are seeking a Senior Systems Software Engineer to join our team, focusing on designing, building, and maintaining large-scale production systems for our Observability and Telemetry Platform. You will ensure high efficiency and availability of our internal and external GPU cloud services, enable developers to make changes to existing systems, and optimize production systems.
Key Responsibilities:
- Design, implement, and support operational and reliability aspects of large-scale Observability & Telemetry collection platforms
- Engage in and improve the whole lifecycle of services, from inception to refinement
- Support services before they go live through system design consulting, software tool development, and capacity management
- Maintain services once they are live by measuring and monitoring availability, latency, and overall system health
- Scale systems sustainably through automation and evolve systems to improve reliability and velocity
- Practice sustainable incident response and blameless postmortems
- Participate in on-call rotation to support production systems
Requirements:
- BS degree in Computer Science or a related technical field
- 5+ years of experience with infrastructure automation, distributed systems design, and large-scale private or public cloud systems
- 5+ years of experience delivering foundational infrastructure and observability platforms
- Experience in one or more of the following programming languages: Python, Go, Perl, or Ruby
- In-depth knowledge of Linux, Networking, and Containers
Preferred Qualifications:
- Interest in crafting, analyzing, and fixing large-scale distributed systems
- Systematic problem-solving approach with strong communication skills
- Experience using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker
- Experience with observability-focused tools like Grafana, OpenTelemetry, and Prometheus
Compensation:
- Base salary range: 152,000 USD - 241,500 USD
- Eligible for equity and benefits
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Systems-Software-Engineer--Observability-and-Telemetry-Platform_JR2024747