# Site Reliability Senior Staff

**Company**: Synopsys
**Location**: Ho Chi Minh City, Ho Chi Minh
**Work arrangement**: onsite
**Experience**: senior
**Job type**: full-time
**Category**: IT
**Industry**: Technology
**Ticker**: SNPS
**Wikidata**: https://www.wikidata.org/wiki/Q2303478

**Apply**: https://careers.synopsys.com/job/ho-chi-minh-city/site-reliability-senior-staff/44408/98572441664?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_a9596e5b-11a

## Description

Synopsys is expanding its Synopsys.ai alongside traditional EDA engineering compute for semiconductor customers in Vietnam. We are looking for an experienced Unix/Linux infrastructure professional to keep these platforms reliable, secure, and operable in Synopsys sites and at major customer environments.

You will join the regional EDA Engineering Compute / platform operations team working with global IT, AE engineering, and product teams. Your work directly enables R&D and customer engineering teams to run EDA and GenAI workloads,from HPC batch grids through containerized AI gateways, vector data services, and agent/MCP-based tool execution on the grid.

The successful candidate will support Synopsys Vietnam engineering compute and data center services, and partner on in-country deployment and sustainment of the related platform components for key semiconductor accounts.

## Responsibilities

### Engineering Compute and Data Center

- Maintain Synopsys Vietnam engineering compute and data center environments per corporate IT, data center, and security standards.

- Support server hardware lifecycle activities: installation, provisioning, maintenance, upgrade planning, retirement, and decommissioning.

- Perform capacity planning, performance troubleshooting, and vendor/customer coordination at customer sites.

- Operate and troubleshoot Linux-based EDA compute: virtualization, engineering resource management, job schedulers, license connectivity, and remote access services.

- Run HPC operations: cluster health, scheduler integration (LSF, Slurm, or similar), InfiniBand where deployed, and performance-related incidents.

- Collaborate with corporate network and information security on data center connectivity, switching, firewalls, routing, and circuits.

### Synopsys.ai Platform Operations

- Support deployment and day-2 operations for Synopsys EDA and AI reference patterns in Vietnam: customer-hosted and air-gapped container environments, API/AI gateways, LLM gateway integration (cloud or on-prem inference endpoints per customer policy), and customer-controlled vector database/GPU compute where applicable.

- Partner on install, upgrade, and configuration of platform components that connect user/agent workflows to the engineering grid,e.g., job services, unified MCP server patterns, and tool execution on LSF/Slurm-backed clusters (in coordination with products and global platform teams).

- Implement and maintain observability, alerting, authentication/authorization integration (e.g., customer IdP, OIDC/SSO/OAuth patterns), and operational documentation for platform services.

- Support agentic AI platform rollouts where components run locally, centrally, or on the grid; escalate design questions to global architecture and engineering teams.

### Reliability, Improvement, and Customer Engagement

- Own or co-own runbooks, monitoring/alerting policies, and automation (scripts, IaC, or equivalent) to reduce toil across the APAC team.

- Apply AI-assisted workflows to improve troubleshooting, knowledge management, and routine operations.

- On-call and incident response: shared rotation (weekday nights, weekends, and holidays); lead or support incident bridges; produce post-incident reports in English.

- Customer air-gapped and on-site support: planned upgrades, break-fix, and customer training for isolated HPC and on-prem AI sites. Travel within Vietnam and APAC as needed (typically 10–20% of time).

- Cross-regional collaboration: handoffs, documentation, and mentor junior staff where applicable.

## Required Skills and Experience

### Technical

- 5+ years in Linux system administration, infrastructure engineering, or production network/storage support.

- Enterprise data center experience: servers, storage, networking, monitoring, and operational processes.

- HPC or EDA engineering compute: Linux clusters, workload schedulers (LSF, Slurm, or similar), and production support.

- Ability to support GPU-enabled compute and containerized platform services in customer-controlled environments.

- Familiarity with virtualization, remote access, observability stacks, and infrastructure automation.

- Understanding of secure multi-tier deployments: gateways, service-to-service auth, and operational telemetry (logs, metrics, tracing) for distributed platform components.

## Benefits

- Comprehensive medical and healthcare plans that work for you and your family.

- In addition to company holidays, we have ETO and FTO Programs.

- Maternity and paternity leave, parenting resources, adoption and surrogacy assistance, and more.

- Purchase Synopsys common stock at a 15% discount, with a 24-month look-back.

- Save for your future with our retirement plans that vary by region and country.

- Competitive salaries.

## Skills

### Required
- Linux system administration
- infrastructure engineering
- production network/storage support
- Enterprise data center experience
- HPC or EDA engineering compute
- Linux clusters
- workload schedulers
- GPU-enabled compute
- containerized platform services
- virtualization
- remote access
- observability stacks
- infrastructure automation

---

Source: [Apply at careers.synopsys.com](https://careers.synopsys.com/job/ho-chi-minh-city/site-reliability-senior-staff/44408/98572441664?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
