Weekday
Description
This position is available on behalf of one of our clients.
Compensation: $60–$90 per hour
A prominent AI research organization is working on the next generation of agent-based evaluation benchmarks for advanced AI models. These sophisticated, multi-stage tasks must be dependable, clearly defined, accurately evaluated, and designed to prevent models from succeeding through shortcuts.
We are looking for experienced QA/Test Engineers to help establish rigorous testing practices and quality standards that ensure each benchmark accurately evaluates the capability it is designed to measure.
Individual tasks may require one to two days of expert-level development and can involve several areas of technical expertise. In this role, you will thoroughly assess these tasks by running them, exploring edge cases, debugging task environments, and uncovering potential weaknesses before they are used in production evaluations.
You will collaborate closely with researchers and task creators in an iterative process focused on identifying issues, improving quality, and validating fixes.
Work Arrangement: Fully Remote — United States
Commitment: Approximately 35 hours per week
Employment Type: Full-Time
Requirements
Key Responsibilities
- Create Comprehensive Test Cases: Build detailed test scenarios to confirm that benchmark tasks behave as expected, covering edge cases, failure scenarios, unusual inputs, and other potential conditions.
- Assess Task Quality: Carefully examine task instructions, requirements, expected results, and reference implementations to uncover ambiguity, contradictions, incomplete requirements, or other quality concerns.
- Debug and Troubleshoot: Apply Python and other development tools to investigate problems in task environments, validation mechanisms, automated tests, and related systems.
- Build QA Methodologies: Develop reusable testing processes, quality checklists, and review frameworks that can be consistently applied across a wide range of benchmark tasks.
- Find Evaluation Weaknesses: Analyze grading systems and AI-agent results to identify loopholes, shortcut opportunities, inconsistent scoring, or other issues that could reduce the reliability of an evaluation.
- Work with Researchers: Partner with researchers and task authors to clearly explain identified problems, suggest practical improvements, and verify that implemented fixes work correctly.
- Uphold Quality Standards: Contribute to the development and enforcement of strong quality standards across sophisticated technical evaluation datasets.
Core Qualifications
- Education: MSc or PhD in a STEM-related field, or equivalent hands-on experience in an engineering-focused or research-intensive environment.
- Experience: At least 1 year of experience in QA, test engineering, software engineering, research engineering, or a similar position involving substantial responsibility for quality.
- Demonstrated ability to create effective test cases, establish QA procedures, and analyze complex technical systems from end to end.
- Practical knowledge of Python and Git, with the ability to navigate unfamiliar codebases, development environments, and technical workflows.
- Strong analytical and debugging abilities, including a systematic approach to locating and resolving technical problems.
- Excellent attention to detail and a consistent approach to maintaining organized, understandable technical documentation.
- Ability to uncover subtle defects, inconsistencies, edge conditions, and unexpected behavior that may not be immediately obvious.
- Experience with AI evaluation, AI training, model testing, benchmark creation, or assessment of AI-generated results is strongly preferred.
- Comfortable working independently on unclear, complex, and open-ended technical challenges.
- Strong written communication skills with the ability to deliver clear and actionable feedback to technical teams.
- A strong quality-oriented mindset, combined with the creativity and persistence needed to discover issues others may overlook.
- Availability to consistently dedicate approximately 35 hours per week.
Ideal Candidate
The ideal candidate brings together strong software testing and debugging expertise with excellent analytical judgment. You should be comfortable challenging complex systems, investigating unexpected outcomes, and determining whether a test actually demonstrates the capability it is intended to measure.
Experience with AI systems, evaluation methodologies, automated grading, or sophisticated technical benchmarks will be especially valuable.
To apply for this job please visit jobs.workable.com.