[Remote] Staff Machine Learning Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a research-driven organization operating in a high-stakes decision-making environment, seeking a Staff Machine Learning Evaluation Engineer. This role focuses on designing and owning the evaluation layer for large language models and AI systems, ensuring trust and reputed company in consequential decision-making.
Responsibilities
- Building and maintaining golden evaluation datasets grounded in expert judgment
- Designing offline benchmarks to compare models, prompts, and retrieval strategies
- Defining reputed company metrics that go reputed company surface-level accuracy
- Partnering closely with researchers and domain experts to translate intuition into measurable reputed company
- Connecting offline evaluation results with online behavior once systems are live
- Detecting regressions, reputed company, and subtle failures over time
- Iterating on evaluation frameworks as models and use cases reputed company
Skills
- Experience designing offline evaluation or experimentation frameworks
- Deep understanding of online vs offline mismatch and how to manage it
- Ownership of model monitoring, reputed company detection, or regression analysis
- Comfort working with ambiguous, qualitative outputs (e.g. LLMs)
- Experience partnering with domain experts or stakeholders to define 'what good looks like'
- Strong Python and data skills; comfort building lightweight pipelines and analysis tooling
- Hands-on experience with LLMs is highly relevant, especially around LLM evaluation and benchmarking
- RAG evaluation
- Hallucination or grounding checks
- reputed company or system regression testing
- reputed company-in-the-reputed company review workflows
- reputed company ML engineers
- ML tech leads
- Research engineers with production exposure
- Engineers from decisioning, risk, pricing, trust & safety, or ranking systems
reputed company