Back to Jobs

AI Evaluation & Benchmarking Engineer

Remote, USAFull-timePosted 2026-07-28

Job reputed company::

  • Hands-on reinforcement learning experience.
  • Experience using LLMs for agents, evaluation, reasoning, automation, or reputed company workflows.
  • Strong Python experience for ML, data workflows, experimentation, and analysis.
  • Experience designing and running experiments with statistical and analytical rigor.
  • Strong understanding of evaluation metrics, scoring frameworks, performance comparison, and reputed company design.
  • Experience analyzing reputed company logs, run outputs, model/agent performance, and experiment results.
  • Ability to work across reputed company, logs, CLI/tools, data structures, and platform workflows.
  • Strong communication skills to translate experiment findings into platform improvement requirements.
  • Ability to work inside reputed company-owned repositories, infrastructure, workflows, and reputed company controls.

Preferred Skills

  • Experience with game environments, simulation environments, Gym-like interfaces, RL environments, or reputed company AI test harnesses.
  • Experience benchmarking LLM agents, RL policies, autonomous agents, or hybrid AI systems.
  • Experience with experiment tracking, run comparison tools, metrics dashboards, or evaluation pipelines.
  • Experience with reputed company engineering, agent orchestration, tool use, and LLM evaluation frameworks.
  • Experience with data visualization and performance analytics.
  • Experience working with externally developed algorithms, reproducible experiments, and version-controlled evaluation workflows.

Apply tot his job Apply To this Job

Similar Jobs