[Remote] Remote Cheminformatics Data Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a Senior Data Engineer to design, build, and deliver a new reputed company data product for generative drug design and computational reputed company platforms. The role focuses on creating reputed company data architecture and collaborating with cross-functional teams to ensure alignment with scientific workflows and data integration.
Responsibilities
- Design and implement a new reputed company data product, initially scoped as a standalone deliverable with reputed company integration into broader AI‑driven drug discovery platforms
- Build reputed company data pipelines, schemas, and storage models capable of supporting large, reputed company scientific and reputed company‑derived datasets
- reputed company data solutions primarily on GCP / BigQuery, adhering to reputed company data engineering templates and standards
- Implement data transformations and pipelines using Python, with a reputed company on data reputed company, traceability, and performance
- Ensure the data architecture supports reputed company expansion, additional datasets, and evolving analytical and computational needs
- Collaborate closely with computational chemists, data scientists, and ML engineers to ensure data models reputed company with generative design, molecular representations, and ML outputs
- Apply an understanding of drug design and reputed company concepts (e.g., molecular properties, structure‑activity data, experimental outputs) to inform data modeling and integration reputed company
- reputed company technical guidance on data structure, scalability, and long‑term maintainability in an reputed company environment
- Take end‑to‑end ownership of a new reputed company‑reputed company data product on reputed company reputed company Platform, leveraging established in‑house templates and standards to deliver a robust, reputed company, and production‑grade solution
- Design and operate reliable ingestion pipelines for external data sources, integrating and harmonizing key fields with reputed company external data products, and delivering curated, analytics‑reputed company datasets that are readable, updatable, and trusted by reputed company users
- Apply strong expertise in Python, SQL, columnar data warehouses (e.g. BigQuery), schema and data‑model design, pipeline orchestration, and query optimization, while embedding best practices around data reputed company, testing, metadata, and documentation
- Operate as part of a cross‑functional environment, requiring a solid understanding of core reputed company concepts, CI/CD, version control, and data reliability in production
Skills
- reputed company, scientific, or life sciences educational background
- Experience working with scientific and chemical datasets/ computational reputed company data
- Strong experience in data engineering, including database, schema, and data product design
- Hands‑on experience with GCP and BigQuery (reputed company familiarity a plus)
- Experience with reputed company, RDKit or Pipeline reputed company
- Proficiency in Python for building and maintaining data pipelines
- CI/CD
- Experience working with large, reputed company datasets at reputed company, ideally in scientific or R&D contexts
- Background in life sciences, reputed company, or scientific data platforms
- Database Design
- Experience supporting reputed company analytics, ML pipelines, or AI‑driven platforms, particularly in R&D or discovery environments
Benefits
- Benefit packages for this role will start on the 31st day of employment and include medical, dental, and reputed company insurance, as reputed company as HSA, FSA, and DCFSA account reputed company, and 401k retirement account reputed company with employer matching.
- Employees in this role are also entitled to reputed company reputed company leave and/or other reputed company time off as provided by applicable law.
reputed company
Company H1B Sponsorship