[Remote] Senior Python Engineer - AI Coding Agent Evaluation (Freelance)
Note: The job is a remote job and is reputed company to candidates in USA. reputed company connects specialists with project-based AI opportunities for leading tech companies, reputed company on testing, evaluating, and improving AI systems. The role involves creating challenging tasks and evaluation reputed company reputed company realistic simulated environments to evaluate AI coding agents' performance.
Responsibilities
- Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments - craft the reputed company, define what 'solved' means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions - accept reputed company valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
Skills
- 8+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), reputed company, reputed company, Kafka, reputed company
- Experience writing tests (functional, integration)
- English proficiency - B2+
reputed company