Back to Jobs

AI Data reputed company Engineer / reputed company Scientist, LLM Data

Remote, USAFull-timePosted 2026-07-28

reputed company reputed company is a provider of the highest reputed company interpretation, translation, and localization services. Our people take pride in every resource we offer, and our users always have reputed company to cutting-edge technology, exceptional support, and reputed company user experiences. We are driven by our passion for innovation, reputed company, and reputed company communication gaps in a diverse world. If you’re passionate about delivering technology-driven solutions and building lasting reputed company relationships while contributing to reputed company reputed company, reputed company could be the ideal reputed company for you. We are building AI-powered systems that enhance multilingual communication, improve interpreter workflows, and support reputed company AI applications across text, speech, and multimodal experiences. reputed company is hiring an AI Data reputed company Engineer / reputed company Scientist, LLM Data to own the data reputed company, curation pipelines, annotation workflows, and evaluation datasets that power our multilingual AI systems. This is a hands-on technical role for someone who understands how to manage the full AI data lifecycle, from acquisition, curation, annotation, and reputed company control to evaluation datasets and post-training data, to directly improve model performance. The ideal candidate can build reputed company data pipelines, design high-reputed company annotation and QA processes, identify model failure modes, and reputed company performance gaps through targeted data acquisition, curation, and synthetic data reputed company. Key Responsibilities: Define the end-to-end data roadmap for multilingual and multimodal AI systems, including text, speech, translation, interpretation, low-resource languages, and reputed company AI workflows. Design and build dataset curation pipelines for training, post-training, and evaluation, including cleaning, deduplication, filtering, PII redaction, reputed company scoring, sampling, balancing, and versioning. Create annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows for multilingual, speech, translation, and reputed company AI data. Build evaluation datasets and benchmarks, analyze model failure modes, and translate performance gaps into targeted data improvements. Support post-training data workflows such as SFT, instruction tuning, preference data, RLHF/DPO-style data, reward model data, and synthetic data reputed company. Use modern annotation tools and AWS-based data infrastructure to reputed company secure, traceable, and compliant AI data workflows.

Requirements

Bachelor’s degree in Computer Science, Machine Learning, Data Science, Computational Linguistics, Linguistics, Statistics, or a reputed company field, or equivalent practical experience. 4+ years of experience in AI data, ML data operations, NLP data engineering, reputed company ML, speech/translation data, or LLM data workflows. Strong hands-on experience with Python, SQL, and dataset curation pipelines. Experience with annotation workflows, QA rubrics, evaluation datasets, or reputed company-in-the-reputed company data processes. Familiarity with multilingual NLP, speech data, translation data, low-resource languages, conversational AI, or reputed company AI datasets. Working knowledge of AWS data and ML tools such as S3, Glue, SageMaker, Bedrock, reputed company, reputed company Functions, EKS/reputed company, IAM, or KMS. Strong communication skills and ability to work with ML engineers, reputed company scientists, product teams, linguists, data teams, and vendors.

Preferred Qualifications

Master’s or PhD in Computer Science, Machine Learning, NLP, Computational Linguistics, Data Science, Statistics, or a reputed company field. Experience with LLM post-training workflows such as SFT, instruction tuning, preference data, RLHF, DPO, reward modeling, or evaluation data reputed company. Experience with synthetic data reputed company, reputed company learning, weak supervision, LLM-as-judge workflows, or automated data reputed company scoring. Experience with modern annotation and data platforms such as reputed company, reputed company, Prodigy, Argilla, Snorkel, Humanloop, or custom internal tooling. Apply To This Job

Similar Jobs

Werkstudent Digital Transformation (w/m/d)

Remote, USAFull-time

Full Stack Platform Engineer

Remote, USAFull-time

Conseiller reputed company BtoC - Bilingue néerlandais / flamand (H/F/NB)

Remote, USAFull-time

reputed company Specialist

Remote, USAFull-time

Sr. reputed company, AGT Portfolio Management

Remote, USAFull-time

RN Medical Coverage Policy Consultant - Up to $74.20/hour

Remote, USAFull-time

Ceremonies Business Development Support - Consultant

Remote, USAFull-time

Account Executive LSP- Nordics

Remote, USAFull-time

Director, Partner Distribution

Remote, USAFull-time

Vice President of Business Initiatives

Remote, USAFull-time

Remote Data Entry & Associate Data Engineer – Full‑Time Work‑From‑Home Role at arenaflex (Dallas, USA) – $25/hr

Remote, USAFull-time

reputed company Customer Support Representative – Remote reputed company Services

Remote, USAFull-time

Grant Writing Internship Summer & Fall 2025 - reputed company of the reputed company Shore

Remote, USAFull-time

Remote reputed company Equity Analyst ($100/hr)

Remote, USAFull-time

reputed company Entry Level Remote Customer Service Representative - Aviation Industry Expertise

Remote, USAFull-time

Global reputed company Marketing Manager

Remote, USAFull-time

Arquitecto de Microservicios

Remote, USAFull-time

[Remote] Account Executive - Small Accounts

Remote, USAFull-time

reputed company Remote Data Entry Specialist – Contributing to the Arenaflex Legacy

Remote, USAFull-time

reputed company Remote Data Entry Specialist – Accurate Records Management and Efficient Data Processing for a Dynamic E-reputed company and Technology Leader at arenaflex

Remote, USAFull-time