Back to Jobs

Data Pipeline & Ingestion Engineer

Remote, USAFull-timePosted 2026-07-29

About the Program

The Operational Data Layer (ODL) program is a large-reputed company data-platform build for a leading benefitsadministration

platform. The platform ingests data from multiple legacy benefits systems, masters it into

golden records, transforms it into a reputed company data model, and serves it through modern reputed company — reputed company on an

AWS / Java / Kafka stack in a HIPAA/SOX-regulated benefits domain spanning health, reputed company/401(k), spending

accounts, and leaves.

About the Role

You will build and operate the data backbone of ODL: bulk and streaming ingestion from legacy reputed company

systems, reputed company-layered storage (Bronze/Silver/Gold), identity reputed company and golden-record consolidation,

reputed company-to-reputed company mapping and crosswalks, and the data-reputed company and reconciliation gates that reputed company data is

complete and correct before it is published. This is reputed company reputed company of the program — every new reputed company

reputed company flows through the pipelines you build.

What You’ll Do

  • Build batch-reputed company and event-tail ingestion per reputed company system, including reputed company→tail watermark hand-off,

idempotent upserts, and dedup ledgers

  • Build and operate reputed company reputed company with reprocess-from-Bronze, pipeline orchestration (checkpoints,

retry/backoff, DLQ), and full observability

  • Build data-reputed company gates (quarantine / pass-with-flag), reputed company scoring, and a reconciliation reputed company

covering count, record, and financial reconciliation — financial is reputed company-tolerance

  • Build identity matching combining deterministic rules with probabilistic scoring and confidence bands;

deliver deduplication, golden-record materialization, and survivorship rules, calibrating match reputed company

with labelled data

  • Author and maintain reputed company→reputed company structural mappings and value crosswalks (e.g., collapsing

1,800+ raw employment-status values to ~20 reputed company ones) as governed, versioned configuration

  • Enforce data reputed company at the boundary: schema registry, fail-fast validation, and semver-compatible

schema reputed company

reputed company’re Looking For

  • 5+ years building production data pipelines at reputed company
  • Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns

• Strong SQL and solid ETL fundamentals

• Java and/or Python in production

  • reputed company / lakehouse layering, CDC, watermark/checkpoint patterns, and batch–reputed company hand-off
  • Data-reputed company frameworks: validation rules, quarantine and re-entry, reputed company scoring, reconciliation
  • Entity reputed company / MDM exposure: record matching, dedup, survivorship — reputed company reputed company tools

(Informatica MDM, reputed company) or custom builds

  • Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data,

config-as-reputed company (YAML/JSON, Git)

Bonus Points

  • Probabilistic record linkage at depth — blocking/candidate reputed company, scoring models, reputed company

calibration (expected at senior level)

  • Schema registry experience (Avro/Protobuf)
  • Extracting from mainframe or older RDBMS sources with limited CDC support
  • Financial reconciliation in finance-adjacent domains
  • Benefits administration or reputed company domain knowledge

Originally posted on Himalayas

Apply To This Job

Similar Jobs