# Site Reliability, Staff / HPC Infrastructure Engineer

**Company**: Synopsys
**Location**: Bengaluru, Karnataka
**Work arrangement**: onsite
**Experience**: staff
**Job type**: full-time
**Category**: IT
**Industry**: Technology
**Ticker**: SNPS
**Wikidata**: https://www.wikidata.org/wiki/Q2303478

**Apply**: https://careers.synopsys.com/job/bengaluru/site-reliability-staff-hpc-infrastructure-engineer/44408/100305008144?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_3b03ee28-2f0

## Description

Synopsys software engineers are key enablers in the world of Electronic Design Automation (EDA), developing and maintaining software used in chip design, verification and manufacturing.

You have spent years keeping large-scale compute environments running, not just available but actually performant, under the kind of load that makes most infrastructure buckle.

### What You'll Be Doing

- Design, build, and optimize large-scale HPC compute farm platforms that support thousands of engineering workloads across global sites

- Administer and tune IBM Spectrum LSF, Slurm, or equivalent workload schedulers to maximize resource utilization and minimize job queue times

- Lead capacity planning and workload optimization efforts, translating engineering demand into infrastructure requirements and deployment timelines

- Drive automation using Python and Shell scripting to eliminate manual toil, improve reliability, and accelerate incident response

- Lead complex troubleshooting and root cause analysis for platform issues, working across storage, networking, LDAP, NFS, and scheduler layers

- Collaborate with R&D, Cloud, Infrastructure, and Security teams on strategic initiatives including cloud-integrated HPC and AI/ML workload enablement

- Mentor junior engineers and provide technical leadership across global teams, setting standards for operational excellence and engineering rigor

- Participate in 24x5 support operations and lead major infrastructure projects from design through deployment

### The Impact You Will Have

- Enable faster product development cycles by ensuring high availability and performance of the compute infrastructure that powers Synopsys EDA tools

- Maximize infrastructure utilization and license efficiency, directly reducing costs and improving engineering productivity across the company

- Reduce incident response time and platform downtime through automation, monitoring, and proactive capacity management

- Accelerate cloud adoption and modernization efforts, helping Synopsys scale compute resources dynamically to meet global engineering demand

- Improve workload throughput and job completion times, giving engineers more iterations per day and faster feedback loops

- Build operational resilience into the platform, ensuring that infrastructure scales reliably as the business grows

- Mentor and elevate the technical capabilities of the global infrastructure engineering team, raising the bar for how we operate and support critical systems

### What You'll Need

- 8+ years of Linux/UNIX systems administration experience with deep expertise in performance tuning, troubleshooting, and large-scale operations

- 5+ years of hands-on HPC or compute farm administration, including workload scheduling, resource management, and capacity planning

- Strong expertise in IBM Spectrum LSF, Slurm, or equivalent schedulers, including policy configuration, job prioritization, and performance optimization

- Advanced knowledge of LDAP, NFS, DNS, enterprise storage systems, and networking in the context of distributed compute environments

- Proven experience with Python and Shell scripting for automation, monitoring, and infrastructure orchestration

- Solid understanding of monitoring and observability tools such as Grafana, Prometheus, Elastic, or Splunk for proactive incident detection and analysis

- Experience with EDA environments is a strong plus, as is familiarity with cloud-integrated HPC platforms on Azure or AWS, Kubernetes, Docker, Ansible, or Terraform

### Rewards and Benefits

- Comprehensive medical and healthcare plans that work for you and your family

- In addition to company holidays, we have ETO and FTO Programs

- Maternity and paternity leave, parenting resources, adoption and surrogacy assistance, and more

- Purchase Synopsys common stock at a 15% discount, with a 24-month look-back

- Save for your future with our retirement plans that vary by region and country

- Competitive salaries

## Skills

### Required
- Linux/UNIX systems administration
- HPC or compute farm administration
- IBM Spectrum LSF
- Slurm
- Python
- Shell scripting
- LDAP
- NFS
- DNS
- enterprise storage systems
- networking
- Grafana
- Prometheus
- Elastic
- Splunk

### Nice to have
- EDA environments
- cloud-integrated HPC platforms
- Azure
- AWS
- Kubernetes
- Docker
- Ansible
- Terraform

---

Source: [Apply at careers.synopsys.com](https://careers.synopsys.com/job/bengaluru/site-reliability-staff-hpc-infrastructure-engineer/44408/100305008144?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
