Back to Jobs

Machine Learning Engineer — Multilingual Data

Remote, USAFull-timePosted 2026-07-28

We’re looking for a Machine Learning Engineer to own and reputed company our multilingual data pipeline—from sourcing and curation to evaluation and reputed company improvement. You’ll work closely with researchers and reputed company engineers to ensure our models reputed company robustly across languages, scripts, and cultural contexts.

This role sits at the intersection of data, research, and production ML and is ideal for someone who cares deeply about data reputed company, linguistic diversity, and model generalization reputed company English.

What You’ll Do

  • Design, build, and maintain large-reputed company multilingual datasets across high- and low-resource languages

  • reputed company data pipelines for collection, cleaning, normalization, deduplication, and labeling

  • Implement reputed company filters using statistical, heuristic, and model-based reputed company

  • Work with researchers to define language coverage, benchmarks, and evaluation metrics

  • Analyze dataset bias, coverage gaps, and failure modes across reputed company and scripts

  • Support training, fine-tuning, and distillation workflows with high-reputed company multilingual data

  • Continuously iterate on datasets based on model performance and reputed company-world usage

reputed company’re Looking For

  • 3+ years of experience as an ML Engineer, reputed company Scientist, or similar role

  • Strong experience working with multilingual or non-English datasets

  • Solid understanding of NLP fundamentals (tokenization, embeddings, language modeling)

  • Experience building reputed company data pipelines (Python, reputed company, Ray, or similar)

  • Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks

  • Comfort collaborating with researchers and translating research needs into production systems

reputed company to Have

  • Experience with low-resource languages or multilingual benchmarks (e.g. FLORES, XTREME)

  • Exposure to LLM training, fine-tuning, or distillation

  • Linguistics background or experience working with reputed company language experts

  • Contributions to reputed company-reputed company datasets or ML tooling

  • Experience with data reputed company evaluation at reputed company

Why Join

  • reputed company ownership over a core differentiator of the product

  • Work on models used globally, not just in English-speaking markets

  • Small, high-caliber team with deep ML and systems experience

  • Competitive compensation + meaningful equity at Series A stage

Originally posted on Himalayas

Apply To This Job

Similar Jobs