Description
At Databricks, we are passionate about enabling data and AI teams to solve the world's toughest problems , from making the next mode of transportation a reality to accelerating the development of medical breakthroughs.
As part of the AI team, you'll build the platforms and products that power everything from data apps, AI agents, model training, model serving, and Vector Search. You'll be joining a high-agency, high-visibility team operating at the frontier of AI infrastructure , with deep ties to research, product, and real-world enterprise use cases.
The Foundation Model Inference team is the backbone of Databricks’ generative AI capabilities. We build the infrastructure that enables our customers to serve, scale, and optimize frontier models with enterprise-grade reliability and performance.
The impact you will have:
- Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)
- Improve reliability, latency, and efficiency of distributed AI workloads
- Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences
- Shape how developers and data scientists build and interact with AI on Databricks
What we look for:
- 8+ years of experience in backend or infrastructure engineering
- Experience with distributed systems, scalable APIs, or cloud-native infrastructure
- Experience with real-time serving, ML infrastructure, or GPU orchestration
- Familiarity with service-oriented architecture, deployment pipelines, and system observability
Bonus points for:
- Exposure to platforms like SageMaker, Vertex AI, or Azure ML
- Contributions to OSS projects like MLflow, PyTorch, Ray, vLLM, SGLang
- Built developer platforms or internal tools supporting AI workflows
The pay range for this role is $190,000-$265,000 USD.