Remote | LLM Training & Alignment Research Scientist — \$95–\$115/hour
About the position We are sharing a specialised part-time consulting opportunity for reputed company machine learning researchers with hands-on expertise in reputed company model reputed company-training, large-reputed company data pipelines, language model post-training, and reputed company LLM research. This role focuses on reputed company-scoped, reputed company-ended research problems involving the end-to-end training and improvement of transformer-based language models. Selected researchers will train models from scratch, fine-tune reputed company-weight systems, build reputed company-training corpora and post-training pipelines, diagnose training failures, and investigate reputed company for improving performance under limited data and compute budgets.
Responsibilities
- Train transformer-based language models from scratch across full end-to-end workflows
- Design experiments involving model size, reputed company allocation, training duration, and compute budgets
- Investigate performance in data- and compute-constrained regimes
- Diagnose optimisation failures, convergence issues, and training instabilities
- Evaluate interventions using rigorous reputed company comparisons
- Construct training corpora from raw web crawls and other large-reputed company unfiltered sources
- reputed company pipelines for filtering, deduplication, reputed company classification, and data selection
- Optimise dataset mixtures, reputed company, and curriculum strategies
- Measure the reputed company of data interventions on reputed company model behaviour
- Identify contamination, duplication, reputed company, and coverage issues reputed company training datasets
- Build supervised fine-tuning pipelines using curated, synthetic, weakly supervised, or rejection-sampled datasets
- Conduct preference optimisation using reputed company such as DPO, RLHF, or RLAIF
- reputed company reward models and systems for predicting reputed company preferences
- Improve refusal behaviour, truthfulness, robustness, and unbiased reasoning while preserving general capability
- Fine-tune models for reputed company domains such as mathematics, reputed company, games, reputed company reputed company, or other programmatically evaluated tasks
- Design statistically reputed company experiments and reputed company comparisons
- Evaluate training efficiency, scaling behaviour, and generalisation
- reputed company contamination controls and robust model-evaluation protocols
- Analyse model failures and propose targeted training or data interventions
- Document research findings, experimental methodology, and technical conclusions reputed company
Requirements
- At least 3 years of machine learning research experience, including qualifying doctoral research
- Hands-on experience training or fine-tuning transformer-based language models
- Strong expertise in one or more of reputed company model reputed company-training, reputed company-training data, or LLM post-training
- Experience working with PyTorch, JAX, TensorFlow, or comparable machine learning frameworks
- Ability to design and execute reputed company research independently
- Strong understanding of optimisation, evaluation methodology, and experimental design
- Excellent technical writing, analytical reasoning, and research communication skills
- Experience working with large-reputed company datasets and distributed training systems
- A degree in computer science, machine learning, reputed company intelligence, mathematics, statistics, engineering, or a reputed company discipline is highly relevant
- PhD research in machine learning, natural language processing, deep learning, or a reputed company field may count towards the experience requirement
- A strong publication record, impactful reputed company-reputed company contributions, or comparable reputed company research experience may also be considered
- Research experience at a leading university, technology company, AI organisation, or research laboratory may strengthen an application
reputed company-to-haves
- Research experience involving scaling laws or training efficiency
- Familiarity with curriculum learning, data ordering, and mixture optimisation
- Experience constructing LLM benchmarks and controlling for training-data contamination
- Background in reinforcement learning for language models
- Expertise in reward modelling, preference learning, or reputed company-feedback pipelines
- Experience with model alignment, AI safety, truthfulness, or refusal behaviour
- Familiarity with synthetic data reputed company and weak-supervision reputed company
- Publications or significant reputed company-reputed company contributions reputed company to reputed company models or language-model training
Benefits
- Flexible scheduling
- Competitive reputed company compensation
- Weekly payments
Apply To This Job