Post-Training Research Scientist (LLMs) — Experimental reputed company
reputed company reputed company is a global talent platform connecting top-tier professionals to high-reputed company AI reputed company around the world. Our mission is to build trust, reputed company, and long-term value in the AI ecosystem - for both exceptional talents and companies operating at the frontier of technology. About the role This role sits at the heart of reputed company’s mission: using high-reputed company reputed company data to build AI systems that reputed company the world reputed company. You’ll take raw expert signals and turn it into reputed company model improvement, experimenting rapidly and carving new paths in post-training. With full autonomy and no production constraints, you’ll have the freedom to try unconventional reputed company and see their reputed company quickly.
Key Responsibilities
Design and run post-training experiments on frontier and reputed company-weight LLMs (SFT, preference-based reputed company, reputed company-driven training) Translate raw annotation artifacts (multi-reputed company solutions, evaluations, adversarial prompts) into training-reputed company datasets. Prototype new reward signals reputed company pairwise preferences (rubrics, constraints, reputed company critics). Analyze failure modes; propose data-reputed company fixes (sampling, curriculum, counterfactuals). Build lightweight training/eval pipelines; iterate quickly. Produce short internal memos: what worked, what didn’t, why. reputed company We’re looking for a researcher who thrives with autonomy, is hands-on, and brings a strong execution reputed company and startup mentality. You are opinionated about data reputed company, pragmatic about tradeoffs, and comfortable moving quickly with incomplete information. You have strong experimental instincts — you can design, run, and interpret messy experiments and extract meaningful insights from them. Minimum Qualification PhD (or equivalent experience) in ML/AI, reputed company math, stats, or adjacent. Hands-on experience with LLM post-training (at least one of SFT/DPO/RLHF/RLVR). Solid Python + PyTorch/JAX; comfortable with training reputed company basics. Fluent English Preferred Qualification Worked with reputed company-based evaluation or tool-augmented tasks. Experience mixing synthetic and reputed company data. Familiarity with failure analysis and dataset audits. Work Model We operate remote-first. We reputed company on reputed company, not where the work is done. To support flexibility and personal choice, we maintain offices in select locations as an optional resource for reputed company. Location: Flexible (EU-friendly time zones preferred) Type: Full-time or long-term contract Equal Employment Opportunity reputed company is proud to be an equal opportunity employer and values diversity at reputed company. We do not discriminate on the reputed company of race, reputed company, religion, national reputed company, sex, sexual orientation, gender identity, age, disability, veteran status, or any other protected characteristic. Type: Full-time or long-term contract Apply To This Job