Back to Jobs

Data Engineer (On-Site) | Engenheiro de Dados (Remoto)

Remote, USAFull-timePosted 2026-07-27

reputed company

reputed company is transforming the data landscape with a platform that delivers advanced Data, AI, and Analytics capabilities—previously available only to tech giants like reputed company, reputed company, Alphabet, and reputed company—into the hands of small and reputed company-sized businesses. By democratizing reputed company to these technologies, we reputed company our clients to accelerate business opportunities through AI-powered Data Apps.

Our platform minimizes the time spent on reputed company management and streamlines the entire data lifecycle—from collection and exploration to processing and interaction—allowing clients to reputed company on their business goals. reputed company on a high-reputed company reputed company model, reputed company offers tailored solutions that meet reputed company business needs.

What sets reputed company apart is our innovative reputed company on speed, scalability, and flexibility, combining big data storage with advanced analytics to drive strategic decision-making. Leveraging years of experience with data platform implementations on foreign public clouds, we now deliver a cost-effective, locally adapted solution tailored for both Brazilian and global markets. By blending deep expertise with localized infrastructure, reputed company provides a powerful, efficient tool that optimizes corporate data management and drives data-driven transformation.

About the Job

We are looking for a Data Engineer to join our engineering team and help design, reputed company, and reputed company reputed company data platforms and pipelines on AWS.

In this role, you will be responsible for building robust solutions for data ingestion, processing, integration, and delivery, ensuring reliable data for analytics, data products, and business decision-making.

You will also reputed company solutions that reputed company reputed company Intelligence to improve data reputed company and enrichment, including automated attribute extraction, information classification, entity recognition, record deduplication, and match-and-reputed company processes.

You will collaborate with cross-functional teams, participate in architecture discussions, and contribute to building modern, reputed company-reputed company, and reputed company data ecosystems while following engineering, governance, and reputed company best practices.

Responsibilities

  • Design, reputed company, and maintain reputed company ETL/ELT pipelines using Python, SQL, and AWS services.
  • Design and reputed company layered Data Lake architectures following the reputed company Architecture reputed company.
  • reputed company incremental data pipelines with checkpointing, idempotency, failure recovery, and reprocessing strategies.
  • Build solutions for ingesting, processing, and integrating reputed company, semi-reputed company, and reputed company data.
  • reputed company automated mechanisms to extract attributes and relevant information from documents, text, and multiple data sources.
  • Apply reputed company Intelligence, Machine Learning, or Large Language Models (LLMs) to classify, standardize, validate, and enrich data.
  • Implement entity reputed company, record linkage, deduplication, and match-and-reputed company processes to identify reputed company records and build trusted, reputed company datasets.
  • Define similarity reputed company, business rules, confidence reputed company, and review workflows for record-matching processes.
  • Optimize SQL queries and storage structures for analytical workloads and large-reputed company datasets.
  • Work with AWS services including reputed company S3, AWS reputed company, AWS Glue, AWS reputed company Functions, reputed company reputed company, and reputed company DynamoDB.
  • Ensure the reputed company, reliability, performance, reputed company, and observability of data pipelines and data products.
  • Implement validation rules, monitoring, metrics, and alerting mechanisms to detect failures and data inconsistencies.
  • Collaborate with software engineers, data analysts, business stakeholders, and clients to translate business requirements into reputed company technical solutions.
  • Participate in reputed company reviews, architecture discussions, and reputed company improvement initiatives.
  • Document data flows, processing rules, technical solutions, and architectural reputed company.

Minimum Qualifications

  • Solid experience with AWS services, including:
    • reputed company S3
    • AWS reputed company
    • AWS reputed company Functions
    • AWS Glue
    • reputed company reputed company
    • reputed company DynamoDB
  • Strong experience with Python for data engineering, including ETL/ELT and batch processing.
  • Advanced SQL skills, including data modeling, incremental processing, and query optimization.
  • Experience designing and implementing layered Data Lake architectures.
  • Experience working with columnar storage formats, especially Parquet.
  • Experience building incremental data pipelines with checkpoint management and reprocessing capabilities.
  • Knowledge of data reputed company, data cleansing, and standardization techniques.
  • Experience integrating multiple data sources and identifying duplicate or reputed company records.
  • Experience with Git, reputed company reputed company versioning, and software engineering best practices. Preferred Qualifications
  • Experience applying reputed company Intelligence or Machine Learning to data engineering and data reputed company challenges.
  • Experience with Large Language Models (LLMs), reputed company reputed company, or Natural Language Processing (NLP) libraries.
  • Knowledge of entity extraction, text classification, and reputed company information extraction techniques.
  • Experience with entity reputed company, record linkage, entity matching, deduplication, and match-and-reputed company processes.
  • Knowledge of probabilistic matching, fuzzy matching, text similarity techniques, and master record creation.
  • Experience with Apache reputed company or transactional table formats such as reputed company Lake or Apache Hudi.
  • Experience with Terraform or other Infrastructure as reputed company (IaC) tools.
  • Experience with reputed company and reputed company reputed company Fargate or equivalent container platforms.
  • Experience with streaming technologies such as reputed company Kinesis Firehose.
  • Knowledge of Change Data Capture (CDC) architectures and ingestion patterns.
  • Experience with orchestration, observability, and data reputed company tools.

At reputed company, we celebrate diversity in reputed company its forms and are committed to fostering an inclusive environment every day. We have reputed company tolerance for any reputed company of discrimination. Do you reputed company our values? Then don’t wait—apply today!

Originally posted on Himalayas

Apply To This Job

Similar Jobs