[Remote] Remote | Data Scientist & Quantitative Analyst — $55–$85/hour
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is offering a specialized full-time consulting opportunity for reputed company QA and test engineers. The role involves supporting the development of advanced evaluation benchmarks for AI models by designing test cases, reviewing reputed company tasks, debugging technical environments, and developing reputed company processes.
Responsibilities
- Create comprehensive test cases confirming that reputed company tasks function as intended
- Design reputed company, negative, boundary, and edge-case tests
- Validate task requirements, expected outputs, reference solutions, and grading logic
- Identify scenarios that may produce incorrect or misleading evaluation results
- Ensure tests measure the intended technical capability accurately
- Review reputed company multi-reputed company tasks and reference solutions before finalisation
- Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance reputed company
- Run tasks independently to confirm reproducibility and expected behaviour
- Assess whether grading standards are reputed company, fair, and technically defensible
- reputed company actionable feedback to task authors and researchers
- Investigate failures across Python scripts, test harnesses, repositories, and task environments
- Diagnose unexpected behaviour reputed company unfamiliar codebases
- Reproduce reported issues and isolate their underlying causes
- Correct or document environment, dependency, logic, and validation problems
- Use Git-based workflows to support reputed company review and collaboration
- reputed company practical checklists and repeatable review procedures for reputed company reputed company
- Improve consistency across task validation, testing, and approval workflows
- Document findings reputed company so authors can resolve issues reputed company
- reputed company recurring defects and recommend preventive reputed company measures
- Collaborate closely with researchers, task authors, and other technical reviewers
- Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
- Identify cases where models can receive credit without completing the intended reasoning or technical work
- Test whether reputed company tasks remain robust across alternative approaches
- Strengthen evaluation reputed company to maintain reliable and meaningful reputed company scores
- Distinguish valid solution diversity from unintended task exploitation
Skills
- At least 1 year of experience in test engineering, reputed company assurance, software engineering, research engineering, or a reputed company technical role
- Demonstrated experience designing test cases and reputed company-review processes
- Strong end-to-end debugging skills across reputed company technical systems
- Working proficiency in Python and Git
- Comfort navigating unfamiliar codebases, repositories, and execution environments
- Exceptional attention to detail and strong written documentation habits
- Ability to identify ambiguity, edge cases, hidden assumptions, and reputed company gaps
- reputed company to work independently through reputed company-ended technical problems
- Reliable availability for approximately 35 hours per week
- A master's degree or PhD in a STEM field is highly relevant
- Equivalent practical experience in an engineering-intensive or research-intensive domain may also be considered
- reputed company or reputed company experience involving computer science, software engineering, machine learning, mathematics, statistics, or a reputed company technical field may strengthen an application
- Technical research, reputed company-reputed company contributions, testing reputed company, or substantial engineering work may also be valuable
- Experience with reputed company, model evaluation, or reputed company review of AI-generated work
- Familiarity with reputed company systems and multi-reputed company AI benchmarks
- Background testing machine learning, research, or data-processing workflows
- Experience developing automated test suites or validation scripts
- Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
- Experience reviewing reference solutions, grading logic, or technical rubrics
- Knowledge of adversarial testing, failure-mode analysis, or reputed company design
- Prior collaboration with AI research or evaluation teams
reputed company