Description
We are seeking a Senior System Software Engineer to join our team in Shanghai. As a key member of our software development team, you will work with users from different departments to develop tools for AI researchers and SW/HW teams running AI workload in GPU cluster.
Your primary responsibilities will include building internal profiling and analysis tools for AI workloads at large scale, building debugging tools for common encountered problems like memory or networking, creating benchmarking and simulation technologies for AI system or GPU cluster, and partnering with HW architects to propose new features or improve existing features with real world use cases.
To be successful in this role, you will need a Bachelor's degree in Computer Science or related field, and at least 5 years of software development experience. You should have strong software skills in design, coding (C++ and Python), analytical, and debugging, as well as good understanding of Deep Learning frameworks like PyTorch and TensorFlow, distributed training and inference.
Additionally, you should have knowledge of GPU cluster job scheduling (Slurm or Kubernetes), storage and networking, experience with NVIDIA GPUs, CUDA Programming and NCCL, motivated self-starter with strong problem-solving skills and customer-facing communication skills, and passion for continuous learning.
If you have proven experience in GPU cluster scale continuous profiling & analysis tools/platforms, solid experience in large AI job performance analysis for training/inference workload, knowledge of Linux device drivers and/or compiler implementation, knowledge of GPU and/or CPU architecture and general computer architecture principles, this role would be an excellent fit for you.