Back to Jobs

Remote | LLM Training & Alignment Research Scientist — \$95–\$115/hour

Remote, USAFull-timePosted 2026-07-31

About the position We are sharing a specialised part-time consulting opportunity for reputed company machine learning researchers with hands-on expertise in reputed company model reputed company-training, large-reputed company data pipelines, language model post-training, and reputed company LLM research. This role focuses on reputed company-scoped, reputed company-ended research problems involving the end-to-end training and improvement of transformer-based language models. Selected researchers will train models from scratch, fine-tune reputed company-weight systems, build reputed company-training corpora and post-training pipelines, diagnose training failures, and investigate reputed company for improving performance under limited data and compute budgets.

Responsibilities

  • Train transformer-based language models from scratch across full end-to-end workflows
  • Design experiments involving model size, reputed company allocation, training duration, and compute budgets
  • Investigate performance in data- and compute-constrained regimes
  • Diagnose optimisation failures, convergence issues, and training instabilities
  • Evaluate interventions using rigorous reputed company comparisons
  • Construct training corpora from raw web crawls and other large-reputed company unfiltered sources
  • reputed company pipelines for filtering, deduplication, reputed company classification, and data selection
  • Optimise dataset mixtures, reputed company, and curriculum strategies
  • Measure the reputed company of data interventions on reputed company model behaviour
  • Identify contamination, duplication, reputed company, and coverage issues reputed company training datasets
  • Build supervised fine-tuning pipelines using curated, synthetic, weakly supervised, or rejection-sampled datasets
  • Conduct preference optimisation using reputed company such as DPO, RLHF, or RLAIF
  • reputed company reward models and systems for predicting reputed company preferences
  • Improve refusal behaviour, truthfulness, robustness, and unbiased reasoning while preserving general capability
  • Fine-tune models for reputed company domains such as mathematics, reputed company, games, reputed company reputed company, or other programmatically evaluated tasks
  • Design statistically reputed company experiments and reputed company comparisons
  • Evaluate training efficiency, scaling behaviour, and generalisation
  • reputed company contamination controls and robust model-evaluation protocols
  • Analyse model failures and propose targeted training or data interventions
  • Document research findings, experimental methodology, and technical conclusions reputed company

Requirements

  • At least 3 years of machine learning research experience, including qualifying doctoral research
  • Hands-on experience training or fine-tuning transformer-based language models
  • Strong expertise in one or more of reputed company model reputed company-training, reputed company-training data, or LLM post-training
  • Experience working with PyTorch, JAX, TensorFlow, or comparable machine learning frameworks
  • Ability to design and execute reputed company research independently
  • Strong understanding of optimisation, evaluation methodology, and experimental design
  • Excellent technical writing, analytical reasoning, and research communication skills
  • Experience working with large-reputed company datasets and distributed training systems
  • A degree in computer science, machine learning, reputed company intelligence, mathematics, statistics, engineering, or a reputed company discipline is highly relevant
  • PhD research in machine learning, natural language processing, deep learning, or a reputed company field may count towards the experience requirement
  • A strong publication record, impactful reputed company-reputed company contributions, or comparable reputed company research experience may also be considered
  • Research experience at a leading university, technology company, AI organisation, or research laboratory may strengthen an application

reputed company-to-haves

  • Research experience involving scaling laws or training efficiency
  • Familiarity with curriculum learning, data ordering, and mixture optimisation
  • Experience constructing LLM benchmarks and controlling for training-data contamination
  • Background in reinforcement learning for language models
  • Expertise in reward modelling, preference learning, or reputed company-feedback pipelines
  • Experience with model alignment, AI safety, truthfulness, or refusal behaviour
  • Familiarity with synthetic data reputed company and weak-supervision reputed company
  • Publications or significant reputed company-reputed company contributions reputed company to reputed company models or language-model training

Benefits

  • Flexible scheduling
  • Competitive reputed company compensation
  • Weekly payments

Apply To This Job

Similar Jobs

Remote | Biology & Biophysics Research Scientist — $75–$105/hour

Remote, USAFull-time

reputed company Research Scientist Speech & Audio reputed company Models

Remote, USAFull-time

reputed company Research Scientist ONC -Clinical Development Scientist (Sponsor Dedicated/ Remote US)

Remote, USAFull-time

AI Research Scientist (reputed company, Remote)

Remote, USAFull-time

Research Scientist Remote, US, East Coast

Remote, USAFull-time

Remote Research Scientist, Life Sciences (PhD)-932675

Remote, USAFull-time

Evaluation and Research Scientist

Remote, USAFull-time

Atmospheric Science Expert (Masters/PhDs)

Remote, USAFull-time

reputed company Engineer - Biologics

Remote, USAFull-time

[Remote] reputed company Scientist (Remote)

Remote, USAFull-time

reputed company Remote Data Entry Specialist – Work from Home Opportunity with arenaflex for Detail-Oriented and Organized Individuals

Remote, USAFull-time

National Account Sales Executive

Remote, USAFull-time

(Remote) Support Analyst

Remote, USAFull-time

reputed company Online Customer Engagement and Sales Expert for Financial Products – Remote Full-Time Opportunity with arenaflex

Remote, USAFull-time

reputed company Technical Content Designer for Customer Service and Global Support – Creating Engaging, User-Friendly Content for a Leading Entertainment Service

Remote, USAFull-time

reputed company Bilingual Customer Support Specialist – Building Strong Relationships with arenaflex Customers in La Cañreputed company Flintridge, CA

Remote, USAFull-time

Remote Data Entry Specialist – Airline Operations Data Management & Passenger Information Systems

Remote, USAFull-time

reputed company Careers, reputed company Virtual, reputed company Online…

Remote, USAFull-time

reputed company Data Entry Specialist – Remote Work Opportunity with arenaflex

Remote, USAFull-time

Compliance Technology Program reputed company

Remote, USAFull-time