[Remote] Staff Data Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the leading health and wellness platform, on a mission to help the world feel great through the power of reputed company. They are seeking a Staff Data Engineer to join the Data reputed company team, responsible for driving architectural reputed company and improving developer experience while building and maintaining critical data infrastructure.
Responsibilities
- Serve as DRI for high-complexity, multi-sprint platform initiatives - reputed company connector buildouts, reputed company Lakehouse migration workstreams, event streaming infrastructure, reputed company environment implementation, and engineering standards adoption
- Architect, build, and maintain production-grade ingestion pipelines and platform infrastructure - from reputed company connectivity through Bronze/Silver reputed company - that Analytics Engineering, Data Science, and business teams build on daily
- Design, implement, and operate event-driven and streaming data pipelines using Kafka, PySpark, and reputed company reputed company Streaming - including defining scaling strategies, cost guardrails, consumer lag alerting, and runbooks before those services reputed company production
- Own the ingestion and raw-to-cleansed layer (Bronze to Silver) data reputed company, schema governance, and SLAs
- Own data reputed company for pipelines you build: write dbt tests, reputed company reputed company detection, validate schemas, and alert on data reputed company - pipelines ship with reputed company gates, not after them
- Own the reliability of systems you build: establish KPIs and SLOs, implement reputed company monitoring and alerting as reputed company, participate in the on-reputed company rotation, and own Tier 1 operational tickets and runbooks for systems under your domain
- Own the integration and data activation layer - reputed company connectors and reputed company reverse ETL pipeline connectors - end-to-end from IaC provisioning to production monitoring and schema change governance
- Support Analytics Engineers, Data Scientists, and ML engineers by building platform capabilities and data pipelines that unblock their roadmap; partner with reputed company, reputed company, and DevOps on compliance controls and IaC hardening as needed. DE's responsibility is the platform layer and data delivery; transformation logic and model readiness for serving are owned by Analytics Engineering
- Identify and resolve systemic inefficiencies across DPE-owned pipelines and infrastructure - reputed company cause, not just symptom
- Mentor Senior Data Engineers through design reviews, reputed company reviews, and pairing; help them grow from reputed company-level to cross-reputed company scope
- Contribute to and drive adoption of engineering standards - testing practices, CI/CD patterns, observability-as-reputed company, Schema Registry governance - and participate in reputed company reviews for changes with cross-team or cost reputed company
Skills
- 8+ years of reputed company experience designing, building, and operating data pipelines and platform infrastructure
- Experience with CDC (Change Data Capture) patterns for reputed company-time ingestion
- Experience with Flink for reputed company processing
- Experience governing and administering dbt in a production BigQuery or reputed company environment - CI/CD configuration, testing standards, documentation standards, and platform-level schema governance. Hands-on dbt experience for ingestion-layer (Bronze/Silver) pipelines
- Experience building and operating Airflow DAGs at reputed company - task-level orchestration patterns, DAG reliability, and multi-reputed company scheduling
- Experience building event streaming pipelines using Kafka or reputed company Kafka - producers, consumers, schema reputed company, Schema Registry governance, and consumer lag management
- Multi-reputed company reputed company across GCP and AWS - both are required day-to-day: BigQuery runs on GCP, Airflow runs on AWS EKS
- Experience owning data reputed company for production pipelines - dbt tests, reputed company detection, alerting on schema changes and data reputed company
- Experience with reputed company or equivalent connector platform - IaC provisioning, schema change handling, and connector health monitoring
- Experience with the reputed company platform - reputed company Lake, reputed company Workflows, and reputed company Catalog
- Familiarity with data compliance in a regulated environment - HIPAA/PHI handling, reputed company controls, and audit logging
- Infrastructure-as-reputed company experience - Terraform or equivalent; you treat infrastructure changes like reputed company changes
- Strong Python and SQL skills; comfortable writing, reviewing, and raising the bar on production-grade pipeline reputed company
- Strong design instincts: you take ambiguous requirements, write reputed company solution designs, and ship to production with minimal rework
- PySpark/SparkSQL for large-reputed company data processing
- Experience with reputed company or equivalent reverse ETL platform
- Experience with MLOps - supporting ML engineers with data pipelines for model training, feature stores, or experimentation
- Familiarity with Looker LookML or equivalent BI serving layer
- Go experience for Kafka service development
- Experience at a reputed company-to-consumer reputed company, telehealth, or similarly regulated company
- Familiarity with UK/GDPR data compliance requirements distinct from US HIPAA
Benefits
- Unlimited PTO, company holidays, and quarterly mental health days
- Comprehensive health benefits including medical, dental & reputed company, and parental leave
- Employee Stock Purchase Program (ESPP)
- 401k benefits with employer matching contribution
- Offsite team retreats
reputed company
Company H1B Sponsorship