New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Dialpad

Sr. Software Engineer, AI / ML Inference Platform

Dialpad
Apply →
senior full-time Buenos Aires

First indexed 7 Sept 2026

Description

Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.

We are hiring a Senior Software Engineer to build the shared AI / ML platform that takes Dialpad’s model-backed capabilities from training through production inference. The AI / ML Platform team builds and operates GPU training infrastructure, model evaluation and lifecycle tooling, and production inference systems running on NVIDIA GPUs in GCP.

Responsibilities

  • Design, build, and improve shared platform capabilities spanning model training, evaluation, artifact management, release, production inference, and operational feedback.
  • Build and operate shared GPU training infrastructure that provides scientists with reliable, reproducible, and efficient environments for model development and experimentation.
  • Improve training-cluster scheduling, workload isolation, capacity management, storage, networking, observability, and accelerator utilization.
  • Develop production-serving pathways for low-latency, high-throughput, and highly available inference workloads.
  • Integrate and adapt model-training frameworks and inference runtimes to meet Dialpad’s requirements for automation, observability, security, and operational control.
  • Improve the performance and efficiency of GPU workloads by reasoning across compute, memory, storage, networking, batching, concurrency, and workload scheduling.
  • Partner with ASR and NLP scientists to translate evolving model capabilities into scalable production designs.
  • Counsel scientific teams on production concerns including reproducibility, evaluation coverage, artifact design, resource requirements, serving feasibility, failure modes, and quality–performance trade-offs.
  • Improve how models and related artifacts are versioned, traced, validated, compared, promoted, deployed, and rolled back across environments.
  • Enable safe releases through representative evaluation, automated quality and performance checks, shadow traffic, staged rollouts, candidate-versus-incumbent comparisons, and fast rollback.
  • Build benchmarking and evaluation infrastructure that measures model quality alongside latency, throughput, saturation behavior, reliability, resource utilization, and cost.
  • Strengthen telemetry, structured logging, tracing, dashboards, alerting, and diagnostic tooling across training and production environments.
  • Use performance data, incidents, developer feedback, and production model behavior to identify and deliver high-value improvements across the AI lifecycle.
  • Reduce recurring manual work by building self-service workflows, clear interfaces, and practical standards that other AI teams can adopt.
  • Lead technical projects from design through production operation, contribute to architectural decisions, and mentor other engineers.

Benefits

Dialpad offers competitive benefits and perks, cutting-edge AI tools, and a robust training program that help you reach your full potential.

Requirements

  • Production engineering experience: Seven or more years of professional software engineering experience, with demonstrated ownership of backend, infrastructure, distributed, or ML platform systems in production.
  • ML systems experience: Experience building or operating systems that support model training, model inference, or the lifecycle connecting them.
  • Strong software fundamentals: Proficiency in Python, Go, or another backend-oriented language, with a record of producing maintainable production software and well-designed interfaces.
  • Cloud and Kubernetes fluency: Hands-on experience with Linux, containers, Kubernetes, cloud infrastructure, CI/CD, deployment automation, and production operations.
  • Accelerated-computing knowledge: Experience operating GPU workloads and reasoning about utilization, memory, storage, networking, scheduling, and workload performance.
  • Training familiarity: Working knowledge of modern model-training workflows, including datasets, experiments, distributed execution, checkpoints, reproducibility, and model artifacts.
  • Applied data-science fluency: An understanding of dataset quality, evaluation design, experimental validity, error analysis, model-quality metrics, and production model behavior sufficient to collaborate effectively with applied scientists.
  • Systems and performance judgment: The ability to find bottlenecks across system boundaries and make reasoned trade-offs among model quality, latency, throughput, reliability, capacity, and cost.
  • Operational judgment: A strong instinct for observability, repeatability, release safety, failure containment, rollback, and whole-system resilience.
  • Technical leadership: The ability to independently lead ambiguous projects, communicate clearly across disciplines, mentor engineers, and influence decisions through sound technical reasoning.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/dialpad/jobs/8785402002