Description
NVIDIA is seeking a Senior HPC Platform Architect to join their HPC Infrastructure team. The successful candidate will be responsible for designing, evaluating, and optimizing the compute infrastructure powering NVIDIA's next-generation silicon design and AI workloads. The salary for this position is not specified.
Key responsibilities include:
- Owning data center architecture reviews for new HPC clusters
- Serving as the BDC representative in cluster build meetings
- Analyzing and validating cluster design choices
- Leading performance benchmarking and profiling of HPC cluster infrastructure
- Driving infrastructure optimization at multiple layers
- Collaborating with platform and operations teams on cluster health and capacity planning
- Evaluating new hardware, storage systems, and networking fabrics
- Continuously improving infrastructure observability and benchmarking frameworks
Requirements include:
- B.E./B.Tech or M.Tech/M.S. with 5+ years of hands-on experience in HPC infrastructure, data center architecture, systems engineering, or a senior SRE/platform engineering role at scale
- Deep understanding of data center architecture fundamentals
- Proven ability to evaluate and challenge infrastructure design decisions
- Experience with OS and kernel-level performance tuning
- Hands-on experience with large-scale Linux HPC cluster administration
- Strong Linux/Unix system administration skills and proficiency in scripting
- Experience running HPC performance benchmarks and profiling tools
- Excellent problem-solving and communication skills
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/India-Bengaluru/Senior-HPC-Platform-Architect_JR2022882