Description
We are looking for a Data Flywheel Infrastructure Engineer to build the infrastructure that continuously turns 1P data, 3P data, model signals, evaluation results, and synthetic data into high-quality training data for frontier LLM and multimodal models.
This role owns the systems connecting Data Acquisition → Governance & Compliance → Curation → Training → Evaluation → Failure Mining → Data Improvement.
A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is secure, policy-compliant, rights-aware, traceable, and auditable.
The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence,ultra-capable systems that remain controllable, safety-aligned, and anchored to human values.
Responsibilities:
- Build 1P & 3P Data Flywheel Infrastructure
- Own Data Governance, Security & Compliance Infrastructure
- Build Policy-Aware Data Acquisition & Curation Systems
- Build Evaluation-to-Data Feedback Loops
- Build Synthetic & AI-Native Data Pipelines
- Build Data Quality, Attribution & Observability
Qualifications:
- Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering
- Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering
- Software Engineering experience using Python, SQL, Spark/Flink/Ray
Preferred Qualifications:
- Experience managing third-party datasets, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions.
- Experience building privacy- and security-aware systems for first-party product or user data, including isolation, access controls, retention/deletion, and purpose limitation.
- Experience with data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration.
- Experience building evaluation → failure mining → data generation → training feedback loops.
- Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization.
- Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories.
- Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data.
- Strong understanding of data governance, security, privacy, provenance, access control, and data lifecycle management.
Data Engineering IC4 – The typical base pay range for this role across the U.S. is USD $119,800 – $234,700 per year.