Description
We are looking for a Data Flywheel Infrastructure Engineer to build the infrastructure that continuously turns 1P data, 3P data, model signals, evaluation results, and synthetic data into high-quality training data for frontier LLM and multimodal models.
This role owns the systems connecting Data Acquisition → Governance & Compliance → Curation → Training → Evaluation → Failure Mining → Data Improvement.
A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is secure, policy-compliant, rights-aware, traceable, and auditable.
Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location.
This role is part of Microsoft AI’s Superintelligence Team, created to push the boundaries of AI toward Humanist Superintelligence,ultra-capable systems that remain controllable, safety-aligned, and anchored to human values.
Responsibilities
- Build 1P & 3P Data Flywheel Infrastructure: scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training.
- Own Data Governance, Security & Compliance Infrastructure: build governance and policy enforcement directly into the data platform.
- Build Policy-Aware Data Acquisition & Curation Systems: develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction.
- Build Evaluation-to-Data Feedback Loops: convert model evaluations and real-world failure signals into actionable data tasks.
- Build Synthetic & AI-Native Data Pipelines: use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation.
- Build Data Quality, Attribution & Observability: develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements.
Qualifications
- Master’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 6+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience.
- Software Engineering experience using Python, SQL, Spark/Flink/Ray.
Preferred Qualifications
- Experience managing third-party datasets, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions.
- Experience building privacy- and security-aware systems for first-party product or user data, including isolation, access controls, retention/deletion, and purpose limitation.
- Experience with data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration.
- Experience building evaluation → failure mining → data generation → training feedback loops.
- Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization.
- Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories.
- Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data.
- Strong understanding of data governance, security, privacy, provenance, access control, and data lifecycle management.
Salary Information
Data Engineering IC5 – The typical base pay range for this role across the U.S. is USD $142,800 – $274,800 per year. Data Engineering IC6 – The typical base pay range for this role across the U.S. is USD $165,600 – $296,400 per year.