Description
About the Role
We are looking for an experienced Engineering Manager – AI Engineering to lead the development of scalable AI platforms and infrastructure while managing high-performing engineering teams. You will drive the design, delivery, and optimization of production-grade AI systems powering AI use cases across Meesho.
Responsibilities
- Lead, mentor, and grow a team of AI engineers , setting technical direction, raising the engineering bar, and owning execution and delivery end to end.
- Architect and scale Meesho's AI platform: cross-region model inference, multi-GPU fleet allocation and management, distributed training, and feature-engineering infrastructure.
- Drive inference optimization across the full stack , GPU kernel tuning, quantization (including outlier/tail-distribution handling), and memory/IO-bandwidth optimization , while building agents that codify and delegate known optimization procedures.
- Optimize open-weight models at both the model and inference-engine level , distillation, quantization, speculative decoding, KV-cache and serving-engine tuning.
- Scale data-science productivity through autonomous, agent-driven workflows spanning feature engineering, model training, and rollout.
- Push the frontier across MLOps, LLMOps, compute efficiency, and distributed ML systems.
- Partner with Product, Data Science, and Platform teams to turn AI capabilities into production impact for millions of users.
- Own the team's operating rhythm: hiring, performance management, sprint planning, and OKRs.
Requirements
- Bachelor's or Master's in Computer Science or a related field.
- 9+ years of software engineering experience, including 2+ years managing engineers.
- Strong hands-on experience with the modern LLM inference stack , TensorRT-LLM, vLLM, SGLang , and with production, low-latency model serving at scale.
- Depth in inference optimization: GPU kernel tuning, quantization, speculative decoding, KV-cache and memory/IO optimization. CUDA / GPU programming experience is a strong plus.
- Experience with distributed training and the frameworks behind it , PyTorch FSDP, DeepSpeed, Megatron, or Ray.
- Experience running GPU fleets in production , Kubernetes (ideally GKE), GPU scheduling and allocation, and multi-region/multi-cluster deployment.
- Familiarity with building LLM-powered agents and agentic workflows, and a point of view on where autonomy can replace manual engineering toil.
- Experience with big-data and streaming stacks , Spark, Flink, or similar.
- Proficiency in Python; systems-level fluency (C++ / Go / Rust) for performance-critical paths.
- Strong leadership, problem-solving, and stakeholder-management skills.
Preferred
- Open-source contributions to inference engines, training frameworks, or ML infra tooling.
- Experience managing GPU cost/efficiency (FinOps) for a large fleet on Cloud and Neo-Clouds.
- Track record building platforms for high-scale consumer products (millions of users).
- Familiarity with observability and reliability for ML systems (SLOs, autoscaling, incident response).
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://jobs.lever.co/meesho/a1cc7ad9-a233-4ba6-b53b-e56116f6ac27