# Staff Engineer, Data Platform (R5659)

**Company**: Shield AI
**Location**: San Diego, California
**Work arrangement**: onsite
**Experience**: staff
**Job type**: full-time
**Salary**: USD 150,000-230,000 per-year-salary
**Category**: Engineering
**Industry**: Technology

**Apply**: https://jobs.lever.co/shieldai/0915a15c-9f24-4851-a375-006bab31d6c6?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply
**Canonical**: https://yubhub.co/jobs/job_a5ac79a0-e44

## Description

Shield AI is seeking a Staff Data Platform Engineer to define and build the data foundation of the AI Factory. The Data Platform provides a unifying, knowledge-graph-centered API layer for human and agentic workflows.

Responsibilities:

- Develop a unifying Graph API: Lead the architecture and implementation of the knowledge graph and multi-modal API layer that serves as the backbone for human, service, and agentic workflows.

- Own DataOps infrastructure: Research, optimize, and maintain the storage, indexing, query, ingestion, and compute infrastructure used throughout the data lifecycle.

- Establish best-practices: Establish durable, best-practice patterns for schema modeling, relationships, lineage, and schema evolution.

- Turbocharge agentic data access: Build APIs that enable agents to retrieve structured, connected, and explainable context rather than relying only on keyword or vector similarity.

- Develop reference architectures: Establish recommended storage and compute profiles, deployment patterns, benchmarks, and operational guidance for both internal and customer-managed infrastructure.

- Advise downstream teams: Partner directly with autonomy, ML, test, infrastructure, product, and customer-facing teams to turn real workflows into reusable platform capabilities from modeling to integrations.

- Build first-party integrations: Deliver integrations that make important data easy to collect and aggregate, including data produced by simulations, test infrastructure, training systems, and edge devices.

- Improve developer experience: Create self-service APIs, SDKs, tools, examples, and diagnostics that make correct data modeling and ingestion the easiest path.

- Drive technical direction: Evaluate emerging data and AI infrastructure technologies, make principled build-versus-buy decisions, and guide implementation across team boundaries.

- Raise operational quality: Establish expectations for observability, performance, reliability, security, data integrity, disaster recovery, and lifecycle management.

Key outcomes:

- Human and agentic workflows use one coherent API for discovering data, traversing relationships, and accessing specialized payloads.

- Teams spend their time deciding how to model and use data rather than repeatedly deciding where and how to store it.

- Data produced at the edge, in simulation, during testing, and in training flows into reusable platform models with minimal integration friction.

- Portable and operational platform capabilities across all deployment environments.

- Downstream teams can adopt the platform through stable APIs and SDKs instead of custom point-to-point integrations.

Required qualifications:

- Significant experience designing and operating distributed data solutions, storage systems, or data-intensive backend services.

- Strong software engineering skills and a record of delivering production systems in languages such as Go and Python.

- Deep understanding of data modeling, API design, schema evolution, identity, consistency, indexing, query planning, and data lifecycle concerns.

- Experience working across multiple storage modalities, such as relational or graph databases, object storage, analytical or columnar systems, and file storage.

- Experience designing reliable ingestion and access paths for high-volume or operationally important data.

- Strong understanding of Kubernetes, Linux, networking, security, storage, observability, and distributed-systems fundamentals.

- Experience deploying data infrastructure across cloud or customer-managed environments using modern Infrastructure as Code and platform engineering practices.

- Ability to evaluate technologies through prototypes, benchmarks, operational requirements, and total lifecycle cost rather than feature lists alone.

- Experience defining architecture and technical standards while remaining hands-on in implementation and debugging.

- Demonstrated ability to collaborate with ML researchers, autonomy engineers, test teams, platform engineers, and product stakeholders.

- Clear technical communication and the ability to make complex data architecture understandable to both specialists and downstream users.

## Skills

### Required
- distributed data solutions
- storage systems
- data-intensive backend services
- software engineering
- Go
- Python
- data modeling
- API design
- schema evolution
- Kubernetes
- Linux
- networking
- security
- storage
- observability
- distributed-systems fundamentals

### Nice to have
- specialized modern databases
- graph-backed retrieval
- agent tooling
- structured RAG
- provenance-aware context construction
- explainable retrieval systems
- S3-compatible APIs
- cloud object storage
- content-addressable storage
- multipart transfer
- large-file lifecycle management
- Apache Arrow
- Parquet
- columnar formats
- time-series data
- high-performance analytical query systems
- OpenAPI
- AsyncAPI
- WebSockets
- generated SDKs
- long-lived public API contracts
- Kubernetes storage and data operators
- Terraform
- Helm
- GitOps
- repeatable platform distribution
- distributed execution technologies
- Ray
- workflow execution
- data lineage
- artifact management
- streaming and ingestion technologies
- Kafka
- NATS
- Redpanda
- event-driven systems
- data and analysis in a robotics or AI domain
- ML data lifecycle systems
- experiment tracking
- dataset management
- evaluation infrastructure
- feature or artifact stores
- model versioning
- observability tools
- distributed tracing
- benchmarking
- data security
- authorization
- governance
- retention
- classification
- auditability across shared platforms

---

Source: [Apply at jobs.lever.co](https://jobs.lever.co/shieldai/0915a15c-9f24-4851-a375-006bab31d6c6?utm_source=yubhub.co&utm_medium=jobs_feed&utm_campaign=apply)
