Description
When 5% of Indian households shop with us, it's essential to build data-backed, resilient systems to manage millions of orders daily. We've achieved this with zero downtime.
At Meesho, you'll design and implement scalable and fault-tolerant data pipelines using frameworks like Apache Spark, Flink, and Kafka. You'll lead the design and development of data platforms and reusable frameworks that serve multiple teams and use cases.
Your responsibilities will include:
- Building and optimizing data models and schemas to support large-scale operational and analytical workloads
- Developing streaming solutions using tools like Apache Flink, Spark Structured Streaming
- Driving initiatives that abstract infrastructure complexity, enabling ML, analytics, and product teams to build faster on the platform
- Ensuring data quality, consistency, and governance through validation frameworks, observability tooling, and access controls
- Optimizing infrastructure for cost, latency, performance, and scalability in modern cloud-native environments
- Mentoring and guiding junior engineers, contributing to architecture reviews, and upholding high engineering standards
- Collaborating cross-functionally with product, ML, and data teams to align technical solutions with business needs
We're looking for someone with 5-8 years of professional experience in software/data engineering with a focus on distributed data systems. You should have strong programming skills in Java, Scala, or Python, and expertise in SQL.
Required skills include:
- 2 years of hands-on experience with big data systems including Apache Kafka, Apache Spark/EMR/Dataproc, Hive, Delta Lake, Presto/Trino, Airflow, and data lineage tools
- Experience implementing and tuning Spark/Delta Lake/Presto at terabyte-scale or beyond
- Strong understanding of Apache Spark internals (Catalyst, Tungsten, shuffle, etc.) with experience customizing or contributing to open-source code
Good to have skills include:
- Contributions to open-source projects in the big data ecosystem
- Hands-on data modeling experience and exposure to end-to-end data pipeline development
- Familiarity with OLAP data cubes and BI/reporting tools
- Working knowledge of tools and technologies like ELK Stack, Redis, and MySQL
We offer competitive compensation, employee-centric benefits, and a supportive work environment. Our holistic wellness program, MeeCare, includes benefits across physical, mental, financial, and social wellness.