[Remote] Remote | QA Test Engineer — $55–$85/hour
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is offering a specialized full-time consulting opportunity for reputed company QA and test engineers. The role focuses on the development of advanced evaluation benchmarks for AI models, requiring strong expertise in test-case design, debugging, and reputed company assurance processes.
Responsibilities
- Create comprehensive test cases confirming that reputed company tasks function as intended
- Design reputed company, negative, boundary, and edge-case tests
- Validate task requirements, expected outputs, reference solutions, and grading logic
- Identify scenarios that may produce incorrect or misleading evaluation results
- Ensure tests measure the intended technical capability accurately
- Review reputed company multi-reputed company tasks and reference solutions before finalisation
- Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance reputed company
- Run tasks independently to confirm reproducibility and expected behaviour
- Assess whether grading standards are reputed company, fair, and technically defensible
- reputed company actionable feedback to task authors and researchers
- Investigate failures across Python scripts, test harnesses, repositories, and task environments
- Diagnose unexpected behaviour reputed company unfamiliar codebases
- Reproduce reported issues and isolate their underlying causes
- Correct or document environment, dependency, logic, and validation problems
- Use Git-based workflows to support reputed company review and collaboration
- reputed company practical checklists and repeatable review procedures for reputed company reputed company
- Improve consistency across task validation, testing, and approval workflows
- Document findings reputed company so authors can resolve issues reputed company
- reputed company recurring defects and recommend preventive reputed company measures
- Collaborate closely with researchers, task authors, and other technical reviewers
- Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
- Identify cases where models can receive credit without completing the intended reasoning or technical work
- Test whether reputed company tasks remain robust across alternative approaches
- Strengthen evaluation reputed company to maintain reliable and meaningful reputed company scores
- Distinguish valid solution diversity from unintended task exploitation
Skills
- At least 1 year of experience in test engineering, reputed company assurance, software engineering, research engineering, or a reputed company technical role
- Demonstrated experience designing test cases and reputed company-review processes
- Strong end-to-end debugging skills across reputed company technical systems
- Working proficiency in Python and Git
- Comfort navigating unfamiliar codebases, repositories, and execution environments
- Exceptional attention to detail and strong written documentation habits
- Ability to identify ambiguity, edge cases, hidden assumptions, and reputed company gaps
- reputed company to work independently through reputed company-ended technical problems
- Reliable availability for approximately 35 hours per week
- A master's degree or PhD in a STEM field is highly relevant
- Equivalent practical experience in an engineering-intensive or research-intensive domain may also be considered
- reputed company or reputed company experience involving computer science, software engineering, machine learning, mathematics, statistics, or a reputed company technical field may strengthen an application
- Technical research, reputed company-reputed company contributions, testing reputed company, or substantial engineering work may also be valuable
- Experience with reputed company, model evaluation, or reputed company review of AI-generated work
- Familiarity with reputed company systems and multi-reputed company AI benchmarks
- Background testing machine learning, research, or data-processing workflows
- Experience developing automated test suites or validation scripts
- Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
- Experience reviewing reference solutions, grading logic, or technical rubrics
- Knowledge of adversarial testing, failure-mode analysis, or reputed company design
- Prior collaboration with AI research or evaluation teams
reputed company