Description
We are seeking a Senior Hardware Architect for our Tegra System-on-Chips (SoC) focused on Reliability, Availability, and Serviceability (RAS). As a Senior SoC Architect, RAS, you will define and drive SoC-level RAS hardware architecture across CPUs, interconnects, memory systems, IOs, safety islands, firmware interfaces, and platform-level components.
You will work with world-class systems architects, RAS experts, design teams, verification teams, validation teams, firmware teams, and software partners to define end-to-end hardware RAS features that improve system resiliency, observability, debuggability, error containment, recovery, and serviceability.
Responsibilities:
- Define and drive SoC-level RAS hardware architecture across CPUs, interconnects, memory systems, IOs, safety islands, firmware interfaces, and platform-level components.
- Own RAS features from concept through architecture specification, micro-architecture alignment, RTL implementation support, design verification, silicon validation, debug, and production readiness.
- Develop architectural requirements for fault detection, correction, containment, isolation, telemetry, error reporting, recovery, graceful degradation, serviceability, and diagnostic observability.
- Work closely with design, verification, validation, firmware, software, and platform teams to ensure RAS features are implementable, verifiable, debuggable, and aligned with system-level requirements.
- Understand the broader SoC architecture and identify how RAS mechanisms interact with performance, power, reset flows, clocks, memory hierarchy, interconnect behavior, firmware-visible controls, and platform software.
- Create hardware specifications, architectural requirements, error-handling flows, design guidance, test plans, and architectural models in SystemC, C/C++, Python, or other relevant modeling environments where applicable.
- Plan and review verification and validation strategies for RAS mechanisms, including error injection, recovery validation, coverage analysis, resiliency modeling, and cross-functional architecture reviews.
- Assist in failure analysis and silicon debug for lab, post-silicon, production, and field findings; develop diagnostic screens and localization methods for latent, intermittent, and environment-sensitive failures.
- Apply RAS architecture principles to high-reliability deployment environments, including space-aware and radiation-aware use cases where single-event effects, memory corruption, logic corruption, or cumulative radiation exposure may impact system reliability.
- Follow industry standards and best practices related to RAS, functional safety, semiconductor reliability, debuggability, verification, validation, and silicon testing.
- Patent novel hardware architecture techniques that improve system resiliency, observability, serviceability, and recovery.
Requirements:
- MS or PhD degree in computer engineering, electrical engineering, or equivalent experience.
- At least 8+ years of SoC architecture, design, verification, reliability, silicon validation, or related hardware development experience.
- Strong understanding of Reliability, Availability, and Serviceability (RAS) in the SoC context, including fault detection, correction, containment, telemetry, recovery, degradation modes, debug visibility, and serviceability mechanisms.
- Experience defining and driving hardware architecture features through the full development lifecycle, including architecture definition, design implementation, verification planning, validation, debug, and production readiness.
- Strong understanding of overall SoC architecture and the ability to reason across micro-architecture, full-chip integration, firmware interfaces, software-visible behavior, platform flows, and customer use cases.
- Meaningful industry expertise in one or more SoC architecture areas such as RAS, safety, debug, clocks, resets, interconnects, memory controllers, IO technologies, platform integration, firmware-visible error handling, or diagnostic infrastructure.
- Hands-on experience with design verification, silicon validation, fault injection, coverage analysis, resiliency modeling, diagnostic development, or reliability validation methodology.
- Familiarity with radiation effects, space operation, or other high-reliability deployment environments is strongly valued, including understanding how hardware architecture can mitigate single-event effects and related reliability risks.
- Excellent analytical, written, and verbal interpersonal skills with the ability to work effectively across architecture, design, verification, firmware, software, validation, and customer-facing teams.
Benefits:
- Equity
- Benefits
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-SoC-Architect--RAS_JR2024446