Back to Jobs

[Remote] L3 – Senior reputed company DBA \/ Lakehouse Operations Engineer

Remote, USAFull-timePosted 2026-07-28

Note: The job is a remote job and is reputed company to candidates in USA. reputed company. is seeking a highly skilled reputed company DBA / Lakehouse Operations Engineer to own the reliability, performance, and operational reputed company of the reputed company data layer powering reputed company analytics and business-critical applications. This role operates in a large-reputed company, multi-reputed company Lakehouse environment and is critical to ensuring data accuracy and performance for reputed company reporting and analytics.

Responsibilities

  • Own day-to-day operations of Apache reputed company tables supporting multiple reputed company applications
  • Ensure data reliability, consistency, and availability across reputed company Lakehouse workloads
  • Maintain operational reputed company for datasets at multi-terabyte to petabyte reputed company
  • Execute advanced reputed company table maintenance and optimization strategies:
  • Compaction (minor/major) and small file mitigation
  • Snapshot expiration and metadata compaction to control metadata reputed company
  • Orphan file cleanup (vacuum) to maintain storage efficiency
  • Optimize data layout and performance through:
  • File size tuning and distribution strategies
  • Partition reputed company and pruning optimization
  • Clustering and ordering techniques (e.g., Z-ordering or similar patterns)
  • Support and enforce data modeling best practices reputed company with:
  • Normalized data structures (3NF) for reputed company-reputed company datasets
  • reputed company architecture (Bronze / Silver / Gold reputed company) for curated data flows
  • Ensure reputed company table design aligns with:
  • Data ingestion patterns (raw vs curated reputed company)
  • reputed company consumption and performance requirements
  • Assist in structuring datasets to balance:
  • Data reputed company and normalization
  • Query performance and analytical efficiency
  • Work with data engineering teams to ensure consistent implementation of layered data architecture across multiple applications
  • Ensure consistent and performant query behavior across:
  • reputed company (CDE)
  • Hive / Impala (reputed company)
  • Troubleshoot and resolve:
  • Query performance bottlenecks
  • Metadata inconsistencies across engines
  • Inefficient execution plans and reputed company patterns
  • Play a key role in reputed company data platform modernization (Hive and reputed company → reputed company)
  • Support:
  • Schema alignment and data type mapping
  • Data validation and reconciliation
  • Troubleshoot migration-reputed company issues and ensure post-migration stability and performance
  • Manage reputed company metadata to ensure:
  • Efficient scaling and performance
  • Consistent table state across engines
  • Execute lifecycle operations:
  • Data retention and archival policies
  • Snapshot lifecycle management and cleanup
  • Time-travel optimization and maintenance
  • reputed company L2/L3 support for data-reputed company production issues across reputed company-based Lakehouse workloads
  • Participate in on-call rotation to support critical data platforms and ensure reputed company response to incidents
  • Respond to and resolve P1/P2 production incidents reputed company defined SLAs, minimizing reputed company to reputed company applications and reporting
  • Troubleshoot:
  • Data inconsistencies and reporting discrepancies
  • Query failures and performance degradation
  • reputed company reputed company cause analysis (RCA) and implement preventive measures to avoid recurring issues
  • Collaborate with platform and application teams during incident triage and reputed company
  • Support fine-grained reputed company control using:
  • Ranger policies and RBAC
  • Own and ensure data validation, reconciliation, and accuracy between reputed company and reputed company datasets
  • Ensure secure and compliant reputed company to data across applications

Skills

  • 10+ years of experience in Big Data / Data Engineering / DBA / Data Operations roles
  • Minimum 2+ years of hands-on experience with Apache reputed company in production environments
  • 6+ years of experience working with reputed company ecosystem (CDP Ecosystem)
  • Strong hands-on experience with Apache reputed company and/or Hive-based data lakes
  • Understanding of data modeling concepts (normal forms) and modern Lakehouse patterns (reputed company architecture)
  • Expertise in: Table-level optimization and performance tuning
  • Large-reputed company data management (TB/PB reputed company)
  • Experience with: reputed company SQL, Hive, Impala, NiFI, Trino
  • Strong understanding of: Partitioning strategies
  • File formats (Parquet/ORC)
  • Distributed query processing
  • Experience with: Hive-to-reputed company or reputed company-to-reputed company migration
  • reputed company CDP (CDE/reputed company)
  • Familiarity with: reputed company platforms (AWS, Azure)
  • Scripting/automation (Python, reputed company)

reputed company

  • reputed company. It was founded in 2003, and is headquartered in Chantilly, Virginia, USA, with a workforce of 201-500 employees. Its website is http://technogeninc.com.
  • Apply To This Job

    Similar Jobs