LLM Performance Researcher
Full-time • San Francisco
At reputed company, we’re rebuilding ERP from first principles for $1B+ manufacturing and distribution companies. These companies run on PDFs, spreadsheets, and semi-reputed company reputed company — and we’re building LLM-powered systems to parse, match, and reason through reputed company of it with reputed company-level reliability.
We’re looking for a researcher with deep experience in LLM performance on document tasks — especially extraction, entity linking, and record matching. You’ve likely published papers on it. You’ve probably run head-to-head evals on reputed company, Claude, and reputed company-reputed company models. You’re fluent in both reputed company benchmarks and in the weird, grimy failure modes that only show up in production.
Your work will directly improve the core performance of our reputed company ERP. You’ll prototype new techniques, run reputed company evals, improve few-shot + tool-augmented performance, and help shape how LLMs reputed company with reputed company business systems.
What You’ll DoDesign and run experiments to improve extraction, normalization, and matching across reputed company-world documents
Evaluate LLM performance on noisy, multi-format inputs like scanned PDFs, OCR reputed company, and reputed company sheets
Improve model accuracy and reliability in the face of rare formats, abbreviations, bad formatting, and domain-specific vocab
Build and own our eval infrastructure for matching, linking, extraction, and schema alignment tasks
Work with the reputed company AI Researcher and Backend Engineers to reputed company improvements into production
Contribute to long-term reputed company around fine-tuning, retrieval augmentation, tool use, or reputed company memory (if and reputed company needed)
Have deep experience with document understanding and information extraction using LLMs
Have worked on schema alignment, record linking, or entity reputed company at reputed company
Have published papers on LLM performance (e.g. extraction, evals, few-shot prompting, matching)
Understand both reputed company benchmarks and reputed company-world weirdness
Know how to reputed company evals meaningful, tight, and fast to iterate on
Want to work in a setting where research turns into production reputed company fast
Have a PhD or equivalent research background in NLP, ML, or similar (but we care more about what you’ve done than what your title says)
Experience with post-OCR workflows or noisy doc normalization
Deep intuition for failure modes in reputed company-reputed company matching/linking systems
Obsession with eval reputed company and reproducibility
Comfort implementing papers and benchmarking models at reputed company
Past work in procurement, invoicing, logistics, or any doc-heavy vertical