Description
As a Sr. SDET in Agentic QA, you will own the test automation and quality frameworks that support Dialpad's AI Voice Agent services.
You will develop automated tests for end-to-end product experiences, from frontend UI to backend services to APIs to audio/text interactions.
Your responsibilities include:
- Owning end-to-end quality for agentic features and workflows, including strategy, development, execution, and release qualification.
- Designing and building automation tooling and frameworks for AI/LLM-driven systems, including prompt flows, agent orchestration, and tool integrations.
- Developing and maintaining evaluation frameworks (evals) to measure response quality, accuracy, and hallucination rates.
- Driving automation coverage (80%+ for critical AI workflows) using deterministic + probabilistic validation approaches.
- Integrating AI quality checks into CI/CD pipelines with fast feedback cycles.
- Building tooling for LLM observability and debugging, including prompt tracing and response analysis.
- Partnering with Applied AI teams on prompt engineering, model selection, and evaluation strategies.
- Designing and executing performance and load tests for AI services (latency, throughput, cost efficiency).
- Identifying and mitigating risks related to hallucinations, bias, safety, and edge cases.
- Defining and tracking AI quality KPIs (task success rates, precision/recall, latency, etc.).
- Participating in design and architecture reviews to ensure systems are testable, observable, and resilient.
- Mentoring engineers and contributing to raising the bar on AI quality engineering practices.
To be successful in this role, you will need:
- 6+ years of experience in software engineering or SDET roles with an emphasis on software development.
- Strong programming skills in Python (preferred), Java, or JavaScript.
- Experience testing distributed, cloud-native SaaS systems and APIs.
- Demonstrated proficiency in coding with AI agents to accelerate development and improve code quality.
- Hands-on exposure to LLMs or AI/ML systems (e.g., OpenAI, Claude, Gemini, or similar platforms).
- Understanding of non-deterministic systems and probabilistic testing approaches.
- Experience building test frameworks and scalable automation systems.
- Familiarity with AI evaluation techniques (benchmarking, golden datasets, human-in-the-loop validation).
- Experience with CI/CD pipelines (e.g., Jenkins, GitHub Actions).
- Strong collaboration skills with the ability to work across distributed teams and time zones.
- Bachelor's degree in Computer Science or equivalent practical experience.
For exceptional talent based in British Columbia, Canada, the target base salary range for this position is $150,500-$175,250 CAD.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/dialpad/jobs/8742092002