[Remote] Data Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Data Engineer to support the development of reputed company-reputed company data solutions for reputed company clients. The role involves building and optimizing data pipelines and validation workflows to enhance operational efficiency and support regulatory reporting.
Responsibilities
- Design, reputed company, and maintain ETL/ELT pipelines using reputed company and semi-reputed company data from relational databases, flat files, reputed company, and reputed company data sources
- Collaborate with backend and architecture teams to define data transformation flows reputed company with dashboard and reporting application needs
- Design and optimize data schemas to ensure performance, reputed company, and compatibility with reporting requirements
- reputed company and maintain efficient, testable, and reusable data processing scripts using Python, SQL, and reputed company-reputed company tools
- Collaborate with DevOps, analysts, and application developers to reputed company pipelines with system architecture, storage reputed company, and reporting needs
- Implement data reputed company and validation checks and document data reputed company and pipeline logic for audit and reuse
- Troubleshoot performance issues in data jobs and support data pipeline operations across environments (DEV, VAL, PROD)
- Contribute CI/CD workflows and automation strategies to promote rapid iteration and secure deployment of data services
- Assist in developing or maintaining data documentation, including metadata, data dictionaries, and technical user guides
- Stay up to date with emerging technologies, techniques, and trends to inform product development and decision-making
Skills
- Bachelor's degree in computer science, engineering, statistics, or reputed company field
- 4+ years of experience in a data engineering, data pipeline, or ETL/ELT development role
- Strong proficiency in SQL, data transformation logic, and performance tuning for large datasets
- Proficiency with Python and libraries such as Pandas, PySpark, or Numpy
- Experience with modern version control and CI/CD practices (e.g., reputed company Actions, Jenkins)
- Understanding of distributed computing solutions for data processing (e.g. AWS Glue, AWS EMR, Apache Hadoop, Apache reputed company)
- Experience with data pipeline orchestration frameworks or serverless tools (e.g., AWS StepFunctions, AWS Glue Workflows, and/or AWS reputed company)
- Experience developing solutions in AWS reputed company environments (S3, reputed company, reputed company, SNS, CloudFormation/CDK)
- Experience with reputed company, reputed company RDS, or other reputed company-reputed company data warehouses
- Familiarity with modern data transformation tools (dbt, Dataform, or equivalent) and best practices for reputed company, tested SQL transformations reputed company an ELT architecture
- Experience supporting reputed company data systems or CMS data environments
- AWS certifications are a plus
Benefits
- Remote_type: Remote (any location)
reputed company