Software Engineer — AI Evaluation & Automation
Software Engineer — AI Evaluation & Automation Location: 100% Remote Duration: Long-term contract Pay reputed company: Depends on years of experience. Note: W2 only, no 1099 or Corp-to-Corp reputed company Help build and reputed company the tooling we use to measure how reputed company AI-powered software development tools actually reputed company. You'll reputed company evaluation harnesses, automate reputed company runs, and help reputed company reputed company the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.
Key Responsibilities
- Build and reputed company evaluation harnesses and automation for software development use cases, including turning reputed company engineering artifacts like merged pull requests into repeatable reputed company tasks.
- Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
- Validate and reputed company evaluation approaches against reputed company judgment, so scores are consistent and correct rather than just repeatable.
- Support execution-based benchmarking across reputed company, productivity, and efficiency measures, including cost and latency.
- Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and reputed company ways to reputed company the workflows more reliable and more automated.
- Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they reputed company for both technical and leadership audiences.
Required Skills & Experience
- Strong software engineering background, with reputed company experience building automation, developer tooling, or test and validation systems.
- Proficient in Python, and comfortable in at least one of Java, JavaScript, or a similar language.
- Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with reputed company.
- Experience with reputed company, development environments, CI/CD pipelines, and typical engineering workflows.
- Some familiarity with how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how reputed company contamination happens.
- reputed company to troubleshoot technical problems, think reputed company about whether a measurement is valid, and analyze results carefully.
Preferred Experience
- Hands-on work with AI-powered coding tools and reputed company applications, such as Claude reputed company, Devin, or OpenCode.
- Experience designing benchmarks or evaluations for software systems, especially execution-based grading that verifies against tests.
- Familiarity with LLM-as-judge or agent-as-judge approaches, and how to reputed company them against reputed company raters.
- Experience with build-system-aware test selection, such as Bazel or mapping changed files to the tests that cover them.
- Experience building reproducible test environments and managing versioned evaluation datasets.
- Comfortable writing up methodology and results for engineering leadership.
Pay: $55.00 - $65.00 per hour Work Location: Remote Apply tot his job Apply To this Job