Description
We are currently seeking an experienced professional to join our team in the role of Associate Director, Software Engineering (Guardrail Platform Control Plane & Observability Track).
Principal responsibilities:
- Design and build the Guardrail Control Plane to manage AI safety policies, detector configurations, evaluation workflows, and enforcement strategies across enterprise AI applications.
- Develop scalable and extensible guardrail platform capabilities including policy lifecycle management, detector registry, configuration management, versioning, rollout/rollback, and governance workflow, and runtime decision management.
- Design and implement system architecture and APIs to enable seamless integration between Guardrail Control Plane, Runtime Engine, AI Gateway, RAG, Agent, and other AI platform components.
- Build AI platform observability capabilities providing end-to-end visibility across AI workloads, including request lifecycle tracking, safety decisions, model interactions, performances metrics, cost insight and operational analytics.
- Design and implement LLM tracing solutions using OpenTelemetry / OpenInference to capture AI execution flows including prompts, models calls, retrieval, embedding, tool calls, guardrail checks, and final responses.
- Develop AI safety evaluation and improvement frameworks including datasets management, regression testing pipeline, LLM-as-a-judge evaluation, human feedback workflows, and production quality monitoring.
- Own hands-on engineering delivery including system design, coding, testing, troubleshooting, performance optimization, and production operation of scalable Guardrail Service.
Requirements:
- Bachelor’s degree or above in Computer Science or related discipline, with 8+ years of professional software engineering experience with strong hands-on experience building production-grade backend platforms and distributed systems.
- Proven experience designing and implementing platform systems such as API platforms, workflow platforms, governance platforms, or internal developer platforms.
- Strong system design skills with experience in service architecture, domain modeling, API design, configuration-driven systems, version management, and extensibility patterns.
- Strong hands-on programming experience with Python (FastAPI) and/or Go, with ability to independently design, implement, and operate backend services.
- Experience building configuration management systems, policy engines, workflow systems, rule engines, or platform control planes.
- Experience designing observability platforms including distributed tracing, telemetry pipelines, metrics, dashboards, alerting systems and operational analytics.
- Hands-on experience with LLM observability concepts including prompt tracing, model tracing, retrieval tracing, tool-calls tracing, agent workflow tracing, and AI quality monitoring.
- Experience with OpenTelemetry, OpenInference, or similar tracing frameworks.
- Experience integrating AI observability platforms such as Arize Phoenix, Langfuse, LangSmith, MLflow, or similar tools.
- Experience with Kubernetes/Docker and at least one public cloud platform (AWS/GCP/Azure).
Tech Stack: Python / Go, FastAPI Kubernetes / Docker Redis / PostgreSQL Kafka / PubSub WebSocket / SSE / gRPC OpenTelemetry GCP / AWS / Azure
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://portal.careers.hsbc.com/careers/job/563774612000774