New The Skills of Tomorrow: how AI-exposed is every skill in 2026? See the data →
Dialpad

AI Systems Engineer

Dialpad
Apply →
senior full-time Buenos Aires

First indexed 14 Jul 2026

Description

We are hiring AI Inference Platform Engineers to build production systems that serve our in-house AI models at scale. This role sits at the intersection of model development, high-performance runtime systems, and cloud infrastructure.

You will help turn trained models and emerging AI capabilities into reliable, observable, low-latency production services running on NVIDIA GPUs in GCP. This is an implementation-heavy systems engineering role focused on model serving, runtime optimization, GPU utilization, deployment safety, traffic management, benchmarking, and production reliability.

Our mission is to shorten the path from promising model capability to dependable production impact. We build the shared infrastructure, standards, and release pathways that allow models to move from candidate artifacts into scalable, rollback-safe inference services with clear performance, reliability, and cost characteristics.

Responsibilities:

  • Design, build, and improve systems that connect AI capability development to production inference
  • Build and improve model-serving pathways for low-latency, high-throughput, high-availability inference workloads
  • Operate and optimize containerized workloads on Kubernetes/GCP, with a focus on efficient use of NVIDIA GPUs, memory, storage, and networking
  • Work with model-serving frameworks and runtimes such as vLLM, Triton, TGI, or similar systems, adapting them to internal deployment, observability, and release requirements
  • Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms for model-backed services
  • Build tooling to measure latency, throughput, cost, saturation behavior, and reliability under realistic production traffic
  • Improve how model and capability artifacts are packaged, versioned, promoted, deployed, and rolled back across environments
  • Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting so engineers can understand model-serving behavior in production
  • Contribute to strategies that improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform

Requirements:

  • 6+ years of professional software engineering experience
  • Production engineering experience with a track record of shipping backend services, infrastructure systems, or production platforms
  • Strong software fundamentals with proficiency in writing maintainable production code in Python, Go, or another backend-oriented language
  • Experience building, operating, or optimizing high-throughput services, distributed systems, data/ML infrastructure, or runtime platforms
  • Kubernetes and Linux fluency with hands-on experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations
  • Strong instinct for reproducibility, observability, rollout safety, failure modes, and whole-system resilience
  • Comfort reasoning about bottlenecks across compute, memory, network, storage, batching, concurrency, and service-level objectives
  • Ability to work closely with model developers, product engineers, infrastructure teams, and technical leadership

Benefits:

  • Competitive salary
  • Comprehensive benefits
  • Real opportunities for growth
  • Cutting-edge AI tools
  • Robust training program
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting: https://job-boards.greenhouse.io/dialpad/jobs/8631737002