Description
Microsoft Ads is looking for a hands-on Principal Software Engineer to build the next generation of agentic evaluation, audit, and human-review platforms.
You will create systems that help teams understand how decisions were made, identify where they fail, investigate supporting evidence, prioritize uncertain or high-impact cases, and turn review outcomes into measurable improvements.
The platform will connect: production decisions → automated evaluation → agentic investigation → human judgment → system improvement
Responsibilities
- Architect and build scalable platforms for agentic evaluation, production audits, and human-in-the-loop review.
- Develop workflows in which AI agents gather evidence, analyze cases, compare decisions, and assist reviewers while preserving human oversight and accountability.
- Build reviewer experiences that present the right evidence, explanations, uncertainty, and recommended actions at the right time.
- Design distributed orchestration systems for long-running workflows involving models, agents, tools, automated checks, and human decisions.
- Provide technical leadership across engineering, applied science, product, design, operations, and policy teams.
Qualifications
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, with 8+ years of relevant industry experience.
- Strong production-coding skills in at least one of C++, C#, Java, Python, or TypeScript.
- Demonstrated experience designing, building, and operating large-scale distributed services or software platforms.
- Strong system-design skills across APIs, workflow orchestration, event-driven systems, storage, concurrency, fault tolerance, and service reliability.
- Experience building AI evaluation, agent orchestration, human-in-the-loop, audit, experimentation, or LLMOps platforms.
- Experience owning substantial systems from initial architecture through rollout, adoption, and production operations.
Preferred Qualifications Background in advertising, trust and safety, fraud, responsible AI, search, recommendations, or another high-scale complex decision domain.