Description
Anthropic's Safeguards team builds systems to detect and mitigate misuse of AI models. We're seeking a Machine Learning Infrastructure Engineer to own the infrastructure behind this research, focusing on scaling, reliability, and efficiency.
Key responsibilities:
- Build and scale infrastructure and data pipelines for Safeguards machine learning research
- Own training, evaluation, and scoring workflows, optimizing the time between idea and result
- Design tooling and interfaces for researchers to use directly
- Implement correctness and sanity checking to ensure trustworthy results
- Transition high-value research workflows from experiments to production-grade jobs
- Improve throughput, cost, and reliability of large-scale inference and scoring workloads
- Collaborate with researchers and engineers to understand and anticipate workflow needs
Requirements:
- Strong software engineering fundamentals and hands-on coding ability in Python
- Experience with data-intensive or distributed systems in production
- Experience building tooling or infrastructure for engineers or researchers
- Comfort working across the research-to-deployment pipeline
- Ability to debug performance and correctness problems
- Strong written and verbal communication skills
Preferred qualifications:
- Experience with high-performance, large-scale machine learning systems
- Familiarity with language modeling and transformers
- Experience with machine learning framework internals, GPU programming, or inference optimization
- Experience building experiment tracking, caching layers, or evaluation harnesses
- Interest in AI system misuse risks and mitigation
The annual compensation range for this role is $350,000-$500,000 USD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/anthropic/jobs/5364804008