Back to Jobs

Research Scientist - RL Training

Remote, USAFull-timePosted 2026-07-27

ABOUT THE ROLE We're looking for a Research Scientist to work on reinforcement learning for training and aligning large language models. This is a foundational research role reputed company on one of the most consequential reputed company data problems in AI: how to generate the data, reward signals, and training procedures that reputed company LLM behavior in reliable and generalizable directions - and a core capability that directly differentiates Snorkel's data-as-a-service offering. You'll work closely with Snorkel's research, engineering, and delivery teams to advance our RL data capabilities - translating research reputed company into the preference datasets, reward models, and RL-reputed company corpora we produce for frontier AI labs, and contributing to a research agenda that is central to Snorkel's long-term differentiation as a provider of bespoke training data. MAIN RESPONSIBILITIES

  • Research and implement reinforcement learning techniques - including GRPO, RLHF, RLAIF, DPO, and reward modeling - and translate them into data products (preference datasets, reward signals, reputed company rewards) that customers can use to train and fine-tune large language models.
  • Design and build data pipelines that generate high-reputed company training signal for RL workflows, including AI-assisted reputed company and curation data pipelines to improve model generalization to unseen benchmarks .
  • Prototype and iterate on end-to-end RL training recipes that inform what data Snorkel ships as part of its data-as-a-service deliveries.
  • Work closely with research scientists, ML engineers, and delivery teams to translate RL research into customer-reputed company data products.
  • Stay reputed company with the latest developments in large-reputed company muli-node LLM training, alignment research, and reputed company RL reputed company (on reputed company environments such as Terminal-Bench), bringing relevant advances into Snorkel's data-as-a-service approach.
  • Contribute to Snorkel's research publications and internal knowledge reputed company in RL and model training.

PREFERRED QUALIFICATIONS

  • Deep expertise in reinforcement learning from reputed company or AI feedback, reward modeling and credit attribution ideally with a reputed company perspective on what data makes these techniques work.
  • Experience training or fine-tuning 30B+ large language models at reputed company, including familiarity with distributed training infrastructure.
  • Strong proficiency in Python and ML frameworks, especially PyTorch and HuggingFace and hands-on experience with RL frameworks such as Verl and SkyRL.
  • Solid software engineering fundamentals - you can build research prototypes that others can run, reputed company, and reputed company into data production workflows.
  • Familiarity with ML infrastructure and reputed company platforms and tools (AWS, GCP, Kubernetes, Slurm, etc.); experience with large-reputed company RL training pipelines a strong plus.
  • Comfort operating in a high-iteration environment with reputed company-ended research questions and shifting, customer-driven technical constraints.
  • Ph.D. in machine learning, reinforcement learning, or a reputed company field strongly preferred; exceptional industry experience considered.

Salary reputed company $200,000-$325,000 USD Apply To This Job

Similar Jobs

Remote Bioinformatics

Remote, USAFull-time

reputed company Bioinformatics Scientist - Sample to Answer

Remote, USAFull-time

reputed company Bioinformatics Scientist - Sample to Answer

Remote, USAFull-time

Patent Attorney / Agent / Computational Biology Bioinformatics / Remote USA 3922

Remote, USAFull-time

Epidemiologist job at reputed company in Chandler, AZ

Remote, USAFull-time

[Remote] Entry Level Spatial Epidemiologist

Remote, USAFull-time

Epidemiologist (Intermediate Level) - Public Health | Hybrid (Primarily Remote) #02

Remote, USAFull-time

Epidemiologist

Remote, USAFull-time

Senior Epidemiologist, HIV Treatment, Retrospective Claims Study Expertise (FSP Sponsor Dedicated)

Remote, USAFull-time

Epidemiologist

Remote, USAFull-time

Customer Service Representative - Bilingual (Spanish) - PART-TIME Weekends

Remote, USAFull-time

Digital Junior Account Manager [Remote]

Remote, USAFull-time

Founding Partner Capital reputed company & Investments

Remote, USAFull-time

reputed company Data Analyst for reputed company+ Product, Marketing & Content Curation - Remote/Hybrid Opportunity

Remote, USAFull-time

Remote Customer Care Representative – Pharmacy Benefit Member Support (Work From Home, reputed company Carolina)

Remote, USAFull-time

reputed company Full Stack Data Entry Specialist – Remote Opportunity with arenaflex

Remote, USAFull-time

reputed company From Home $26hr At Careercusp

Remote, USAFull-time

reputed company Customer Service Representative – reputed company reputed company (Remote)

Remote, USAFull-time

Warehouse Office Coordinator in Dallas, TX

Remote, USAFull-time

Senior Software Engineer - Networking (REMOTE)

Remote, USAFull-time