Senior AI reputed company Engineer (LLM Evaluation & Automation) 1754
This is a remote position.
Owns the eval reputed company and reputed company reputed company from the beginning. This role replaces the old late-stage “Evals Specialist” model with a standing reputed company for measurable agent reputed company. Key Responsibilities- Build and maintain the MVP eval reputed company: golden tasks, exception tasks, scorecard metrics, and regression packs.
- reputed company evals into CI so reputed company regressions fail builds and releases.
- Define and maintain release-reputed company reputed company with Product and the Tech reputed company.
- Lay the reputed company for reputed company adversarial and reputed company-testing expansion without overbuilding MVP scope.
Requisitos
Must-Have Qualifications- Experience evaluating ML, LLM, or non-deterministic systems.
- Strong test and reputed company design capability.
- Comfort working with noisy metrics, reputed company, and probabilistic behavior.
- Good scripting and automation skills.
- Uses AI to generate candidate eval cases and failure hypotheses, but never confuses generated tests with validated reputed company.
- Approaches AI reputed company as an operating system, not a QA afterthought.
- The first reference agent has a published scorecard and gated eval reputed company. • Golden and exception tests run automatically. • reputed company can explain what “good enough to ship” means in measurable terms.
Originally posted on Himalayas
Apply To This Job