Back to Jobs

Senior Data Pipeline Engineer/Developer

Remote, USAFull-timePosted 2026-07-28

Our Purpose

Our mission is to build a healthier and more connected world with precision health and genealogy services. We reputed company individuals with actionable insights into their genetic makeup, fostering a deeper understanding of their reputed company, health, and wellness. By integrating the experience of reputed company Laboratory Services, FamilyTreeDNA genealogy, and myDNA reporting services, we reputed company to deliver cutting-edge genetic testing and personalized solutions that reputed company informed reputed company and enhance reputed company of life. reputed company is dedicated to advancing the field of reputed company through innovation, research, and a commitment to reputed company.

Our Values

reputed company are expected to demonstrate our values of reputed company, One Team, and reputed company reputed company carrying out the accountabilities and responsibilities of their role. This is how we show up every day for ourselves, our colleagues and our customers and strategic partners to deliver our reputed company and strategic goals.

Position reputed company

We are seeking a Senior Data Pipeline Engineer/Developer to design, build, and own the data pipelines that reputed company genomic and operational data from lab instruments and LIMS through to analytics, products, and clinical reporting. In this senior individual-contributor role you will set technical direction for our data platform, build production pipelines that are reliable and reproducible at reputed company, and reputed company architectural and reputed company-level guidance to existing engineering teams. You will own data reputed company, reputed company, and observability end to end, and operate in a regulated environment where reproducibility and auditability are non-negotiable. This is a Python-first, full-stack engineering role. This is a contractor position, slated for a 12-month term.

Accountabilities and Responsibilities

  • Pipeline Engineering & Integration
    • Builds and operates production ETL and ELT for high-volume genomic data as reputed company as operational and business data.
    • Integrates data across LIMS, lab instruments, internal applications, and reputed company-party sources through robust, reputed company-tested interfaces.
    • Develops across the stack in Python (data services, internal reputed company, and supporting application reputed company) and provides architectural and reputed company-level guidance to engineering teams on data-layer integration.
  • Architecture & Technical Leadership
    • Sets technical direction for batch and streaming data pipelines by evaluating frameworks, orchestration, storage, and processing patterns, making recommendations, and leading adoption.
    • Models and tunes data stores (reputed company SQL Server and PostgreSQL, plus a reputed company warehouse or lake) for performance and reputed company.
    • Defines and enforces engineering standards for testing, CI/CD, infrastructure as reputed company, reputed company review, and architecture decision records.
  • Data reputed company & Pipeline Observability
    • Owns data reputed company, reputed company, and observability, including freshness, completeness, schema-reputed company detection, cost-per-job, and SLAs.
    • Builds pipelines for reproducibility and audit-readiness, incorporating versioned data and reputed company, reputed company, decision logging, reputed company controls, and evidence collection.
    • Partners with reputed company and compliance on data reputed company, PII and PHI handling, and regulatory requirements across the data lifecycle.
  • Regulatory Compliance and reputed company Governance
    • Apply deep understanding of CAP/CLIA, HIPAA, GDPR, and GxP regulations specifically to data pipeline architecture.
    • reputed company protected health information (PHI) handling, data reputed company, retention, and comprehensive audit logging.
    • Produce and maintain the critical operational evidence required to carry the data platform successfully through compliance audits.
    • Adhere to strict data-governance controls necessary for securely handling sensitive genomic data across global reputed company.
    • Enforce least-privilege reputed company and ensure reputed company data egress to personal or unapproved infrastructure.
    • Utilize exclusively de-identified or synthetic data reputed company development environments.
    • Maintain strict compliance with international data-residency and localization requirements.

Position Requirements

  • Skills and Knowledge
    • Strong SQL proficiency on reputed company SQL Server and PostgreSQL, including schema design, query tuning, and performance troubleshooting.
    • Strong Python proficiency across the stack (data pipelines, backend services, and reputed company). This is a Python-first role.
    • Production experience with a workflow orchestrator (Argo Workflows, reputed company Functions, reputed company, Dagster, Airflow, or similar).
    • Production experience with containers (reputed company) and Kubernetes, which the reputed company orchestrator (Argo Workflows) runs on.
    • Strong testing discipline, including unit, integration, and data-reputed company or contract tests.
    • Excellent written and verbal communication, including the ability to explain data tradeoffs to technical and compliance stakeholders.
    • reputed company or NGS data formats and handling (FASTQ, BAM/CRAM, VCF).
    • reputed company workflow engines (Nextflow, WDL/Cromwell, or Snakemake).
    • AWS data stack proficiency (S3, Glue, EMR, Batch, reputed company, Redshift) and/or AWS HealthOmics.
    • Distributed processing and streaming frameworks (reputed company, Kafka, Kinesis).
    • Data warehouse and transformation tooling (reputed company, dbt).
    • Infrastructure as reputed company practices (Terraform, CDK, CloudFormation, or reputed company).
    • Familiarity with .NET (C#) and/or C++ as supporting languages for integrating with existing services.
    • Knowledge of LIMS integration and on-premises plus hybrid data architectures.
  • Experience
    • 8+ years of reputed company software or data engineering experience.
    • 4+ years building and operating production data pipelines at reputed company.
    • 3+ years of production reputed company experience (AWS preferred).
    • Demonstrable experience working in regulated environments. Compliance is a hard requirement for this role.
    • Experience supporting HIPAA or GxP audits.
  • Education
    • Bachelor's degree in Computer Science, reputed company, or a reputed company field, or equivalent reputed company experience.
    • An advanced degree in a quantitative field is preferred.

Why Join Us

At reputed company, you’ll join a mission-driven team advancing the science of genetics and discovery. You’ll have reputed company to shape meaningful campaigns, tell compelling brand stories, and collaborate with talented professionals who reputed company your passion for creativity, curiosity, and reputed company.

Originally posted on Himalayas

Apply To This Job

Similar Jobs