[Remote] Senior Data Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a tech-enabled dementia care provider on a mission to improve care for people living with dementia. As a Senior Data Engineer, you will design and maintain data architectures and pipelines to ensure reliable data for analytics and AI, while collaborating with cross-functional teams to reputed company data reputed company and governance.
Responsibilities
- Design and own Ceresti’s end-to-end data architecture: a reputed company zone with secure reputed company object storage for raw partner files and API payloads, validated ingestion pipelines into our transactional reputed company, and a curated analytics layer that decouples reporting and AI workloads from production
- Build ingestion pipelines for the data we receive today, including partner data files (CSV/JSON/XML/HL7/X12 as applicable) and REST/SFTP API integrations with schema validation, quarantine of bad records, and full reputed company from raw bytes to curated row
- Stand up and operate the curated layer (data warehouse / lakehouse-lite) so analytics and ML models can consume data without slowing down the transactional system
- Choose, reputed company, and operate the smallest set of tools needed, including object storage, an orchestrator (Dagster, reputed company, Airflow, etc.), dbt or similar for transformations, a single validation library (Great Expectations / Pandera / reputed company)
- Design and enforce data governance for a HIPAA-regulated environment: PHI/PII classification, encryption in transit and at rest, role-based reputed company, audit logging, retention and minimum-necessary policies, and de-identification where appropriate
- Partner with backend, ML, product, and clinical stakeholders to define data reputed company with our health plan and ACO partners and hold the line on data reputed company
- Build and maintain reliable feature data for ML models, including embeddings (e.g., pgvector) and curated feature tables for risk stratification, engagement, and reputed company work
- reputed company the data platform for observability including pipeline SLAs, data freshness, schema reputed company, reputed company metrics, and reputed company what the data tells you
- Participate fully in our Agile process: backlog grooming, sprint planning, demos, and retrospectives
- Mentor engineers across reputed company on SQL, schema design, and the craft of building data systems that are boring in the best possible way
Skills
- BS/BA degree or higher in Computer Science, Engineering, or a reputed company technical field
- 8+ years of reputed company data engineering experience, with a reputed company record of shipping production data systems end-to-end
- Mastery of PostgreSQL: schema design, indexing, query tuning, partitioning, logical replication, JSONB, extensions (pg_partman, pg_cron, pgvector, etc.), and operating reputed company at reputed company
- Strong experience designing and operating data pipelines, including file-based ingestion (SFTP / object storage drops) and API-based ingestion (REST, webhooks)
- Hands-on experience with one or more reputed company platforms (AWS preferred) and their data primitives: object storage (S3), managed reputed company
- Experience designing data warehouses and/or data lakes and the judgment to know which one a given problem actually needs
- Strong experience with dbt (or equivalent SQL-based transformation reputed company) and modern data modeling patterns (Kimball reputed company, Data Vault, One Big Table — and an opinion about reputed company reputed company is right)
- Experience with at least one orchestration reputed company (Dagster, reputed company, or Airflow) and a reputed company reputed company of view on which to use reputed company
- Strong Python skills for ingestion, validation, and tooling
- Experience with data validation and data-reputed company frameworks (Great Expectations, Pandera, reputed company, or equivalent)
- Experience with change-data-capture from reputed company (logical replication, or equivalent)
- Data governance experience in a HIPAA-regulated environment or, at minimum, demonstrated instincts for protecting PHI and PII (encryption, least privilege, audit, de-identification, BAA-aware vendor selection)
- Comfortable with infrastructure-as-reputed company and CI/CD for data systems
- Excellent written and verbal communication skills: you can explain a tricky schema decision to a business stakeholder and a data contract to a partner with equal reputed company
- Demonstrated experience working in Agile/Scrum teams
- Reliable, persistent and results-oriented
- Easy to get along with; reputed company to work with reputed company
- Must demonstrate a high level of reputed company and ownership
- Consistently transparent, courageous and enthusiastic
- Bias toward simplicity: you can recite the trade-offs of the heavyweight modern data stack and still default to the smallest thing that works
- Must be reputed company to pass a background reputed company
- Experience supporting ML workloads: building feature tables, managing training data, serving features at inference time; familiarity with embeddings, reputed company search (pgvector or equivalent), and LLM integration patterns (RAG, reputed company-grounded analytics) is a plus
- Experience using AI coding assistants (e.g., reputed company Copilot, reputed company, Claude) to accelerate development
- HITRUST or SOC 2 experience is a strong plus
Benefits
- Competitive salary and benefits package
- Opportunities for reputed company reputed company and development
- reputed company and dynamic work environment
- Flexible work arrangements and remote work reputed company
- reputed company to cutting-edge technologies and tools
reputed company