Description
ZoomInfo is where careers accelerate. We move fast, think boldly, and empower you to do the best work of your life.
We're looking for a Software Engineer III to join the Data Acquisition team - one of ZoomInfo's most strategically important engineering areas. In this role, you will help build the systems that acquire, transform, validate, and store large raw data sets from a wide range of sources.
Responsibilities
- Design and build backend data pipelines that ingest, validate, normalize, enrich, and store high-volume raw data from multiple sources.
- Develop ETL/ELT workflows and processing jobs using technologies such as Apache Airflow, Apache Beam, Google Dataflow, DataProc, Spark, Kafka, or Pub/Sub.
- Implement new features in Java-based services and data processing applications that support data acquisition at scale.
- Work with batch and streaming architectures for scheduled, near-real-time, and event-driven data flows.
- Improve data quality through schema validation, deduplication, enrichment, monitoring, retries, and controlled backfills.
- Contribute to observability for pipeline health, throughput, latency, cost, and error rates.
- Collaborate with product managers, data teams, and platform teams to translate business requirements into reliable technical solutions.
- Help design, plan, and execute the roadmap for next-generation data acquisition technologies.
Must-Have Qualifications
- 3+ years of professional software engineering experience.
- Solid experience building and operating production data pipelines, ETL/ELT workflows, or data processing systems.
- Strong proficiency with Java and object-oriented programming.
- Hands-on experience with data processing and orchestration technologies such as Apache Beam, Apache Airflow, Spark, Google Dataflow, or DataProc.
- Experience with streaming technologies such as Kafka, Google Pub/Sub, or similar systems.
- Understanding of batch processing, streaming processing, data modeling, schema design, and data quality practices.
- Experience building backend services, APIs, or distributed systems that run in production.
- Experience with at least one cloud provider, preferably GCP.
- Familiarity with cloud data and compute services such as BigQuery, GCS, GKE, Dataflow, DataProc, and Pub/Sub.
- Practical knowledge of SQL, large-scale storage/query systems, logging, monitoring, and alerting.
- Ability to troubleshoot pipeline failures, data issues, performance bottlenecks, and operational incidents.
Nice to Have
- Experience with Kubernetes, especially GKE or EKS, for distributed workloads.
- Experience with Snowflake, BigQuery, Starburst/Trino, or similar data warehouses and query engines.
- Experience with Terraform or other infrastructure-as-code tools.
- Exposure to data integration patterns involving CRM systems, email/calendar APIs, third-party feeds, or large external datasets.
- Experience in a B2B data company, data marketplace, or data-as-a-product environment.
Benefits We offer comprehensive benefits, holistic mind, body and lifestyle programs designed for overall well-being.
This listing is enriched and indexed by YubHub. To apply, use the employer's original posting:
https://job-boards.greenhouse.io/zoominfo/jobs/8604692002