# Incident Response Engineer-Facility Operations Center

**Company**: NVIDIA
**Work arrangement**: remote
**Experience**: senior
**Job type**: full-time
**Category**: IT
**Industry**: Technology

**Apply**: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Australia-Remote/Incident-Response-Engineer-Facility-Operations-Center_JR2020253-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_f2a6c77c-12a

## Description

We are seeking a highly motivated and skilled Incident Response Engineer to join our Facility Operations Center (FOC) team. In this critical role, you will be responsible for coordination and presentation within NVIDIA's datacenters, focusing on incident response, vendor support, and maintenance performance.

**Key Responsibilities:**

- Coordinate and communicate across NVIDIA's datacenter portfolio regarding incidents, maintenance, and reporting/monitoring.

- Develop standards and programs to support reliability and operations initiatives, including Problem and Change Control.

- Analyze failure data and work with machine learning and AI teams to predict future failures.

- Coordinate disaster recovery tests, liaise during audits, and collaborate with internal partners.

- Own and present key business metrics related to incident response.

- Assess process improvement opportunities and partner with process owners.

- Work multi-functionally with other team members and groups within the organization.

**Requirements:**

- Bachelor's degree in a related field (e.g., Electrical Engineering, Computer Science).

- 5+ years of operations or environmental, health, and safety experience within data centers.

- Proficient in developing and driving reliability activities.

- Commercial and financial awareness, with a full comprehension of the impact of failure.

- Highly developed numeracy, statistical, and reporting skills.

- Proficient in the use of asset database and DCIM solutions.

- Experience in designing, deploying, or maintaining large-scale datacenter infrastructure.

**Nice to Have:**

- Proven experience in reliability engineering related to electrical or mechanical cooling systems.

- Certifications such as CDCMP, CMRP, CRL, CRE in Maintenance and Reliability.

- Knowledge of relevant ISO standards and their implementation.

- Demonstrated expertise in statistics, forecasting, and management information methods.

## Skills

### Required
- reliability engineering
- incident response
- datacenter operations
- problem control
- change control
- machine learning
- AI
- asset database
- DCIM solutions
- Microsoft Office Suite
- G-Suite software

### Nice to have
- electrical engineering
- mechanical engineering
- cooling systems
- CDCMP
- CMRP
- CRL
- CRE
- ISO standards

---

Source: [Apply at nvidia.wd5.myworkdayjobs.com](https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/Australia-Remote/Incident-Response-Engineer-Facility-Operations-Center_JR2020253-1?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
