Description
We are seeking a highly motivated Senior HPC Support Engineer – Ethernet / AI Infrastructure to engage with one of our prestigious customers onsite and remotely. The successful candidate will be the primary point of contact for this customer, spending a minimum of one week per month at the customer site in College Station, TX, US, supporting technical questions, debugging, and issue resolution.
As a member of our NVEX Global Technical Support team, you will be a conscientious, proficient communicator who takes ownership of resolving issues while maintaining a high level of customer satisfaction. A significant part of the role involves collaborating with Engineering, Marketing, and Support teams regularly on technical issues.
Responsibilities:
- Resolve sophisticated customer concerns and technical issues through meticulous research, reproduction, and problem-solving for customers installing our products and supporting systems using Linux Operating Systems (multi-distro), focusing on NVIDIA Ethernet Switching technologies and our End-to-End Solutions such as NVIDIA Spectrum-X.
- Respond to customer product support inquiries via telephone, email, or conference calls.
- Resolve customer issues during installation, operation, maintenance, or product application or interoperability with other vendors.
- Participate in multi-functional team meetings and provide feedback to engineering and marketing regarding product requirements, customer experience, support tools, etc.
- Develop, redefine, and document standard methodologies to provide to internal teams (Support/R&D) for support processes and improvements.
Requirements:
- 6+ years of experience providing in-depth Customer Support and debugging for hardware and software products.
- An academic degree from an accredited university or college in Networking, Computer Science/Engineering, or Electrical/IT (or equivalent experience).
- Experience with established AI technologies in day-to-day job responsibilities.
- Knowledge of Enterprise platforms and systems engineering, including Linux triage, servers, and resolving hardware and/or OS internal issues.
- Intellectual curiosity, positive attitude, flexibility, analytical ability, self-motivation, and team-oriented with professional-level communication skills, interpersonal skills, and the ability to maintain and lead the overall resolution for any critical issue raised by our customer.
Technical Skills:
- Networking Technology, protocols, and routing, including TCP, UDP, Ethernet, IP, L2, L3 (ARP, STP, LACP, MLAG, IGMP, PIM, BGP, OSPF), on Enterprise Level.
- Linux OS, including System Administration and Networking (LFCS / RHCSA).
- Debugging networking protocols using tools such as TCPDUMP and Wireshark or similar packet generation and analysis tools.
- Deep understanding of at least two of the following: data centers, servers, distributed systems, virtualization, deep learning frameworks, containers/containerization (i.e., Docker, Kubernetes).
- Adoption of AI solutions like Cursor, Gemini, ChatGPT, Copilot, Glean, etc., in your daily work routine.
Preferred Qualifications:
- Experience in solving problems in large-scale networking and AI Infrastructure environments with overlay technologies (BGP, OSPF, VXLAN, EVPN), RoCE, and QoS Concepts.
- Linux, Networking, and NVIDIA AI Infrastructure and Operations Certifications such as CCONP, CCIE, JNCIE-DC/ENT, RHCE, LFCS, NCP-AII/AIO/AIN.
- Shell Scripting (Python, bash, Ansible, yaml, etc…).
- Effective and comprehensive fixing/debugging methodology.
You will also be eligible for equity and benefits.