AI Evaluation Engineer (Python, QA or Security)

Mindrift

Description

Please provide your CV in English and mention your current level of English proficiency.

Mindrift offers specialists project-based opportunities to work with leading technology companies on testing, assessing, and enhancing AI systems. These assignments are project-based and do not represent permanent employment.

About the Opportunity

We are developing a dataset designed to assess AI coding agents and their ability to handle realistic software development challenges.

Your responsibilities will include creating demanding tasks and clear evaluation standards within realistic simulated development environments:

  • Create authentic developer environments — Build a virtual company setting containing a codebase, infrastructure, and supporting context such as tickets, documentation, conversations, and development history.
  • Develop tasks from realistic project states — Create the task prompt, establish clear success criteria, and make sure an AI coding agent can realistically complete the assignment.
  • Develop solution-verification tests — Create tests that recognize valid implementations while identifying incorrect solutions without being unnecessarily restrictive or overly permissive.
  • Refine tasks and tests through QA feedback — Examine agent-generated solutions, investigate failures, and improve tasks and testing criteria until the evaluation is consistent, reliable, and fair.

What This Role Is Not

  • This is not data labeling.
  • This is not prompt engineering.
  • This is not traditional coding from scratch — the AI agent will produce most of the implementation, while you provide guidance, assessment, and evaluation.

What We’re Looking For

  • At least 5 years of professional software development experience
  • Experience with the core technology stack:
    • Python / FastAPI
    • JavaScript / TypeScript / React
    • Docker
    • PostgreSQL
    • Kafka
    • Redis
  • Practical experience creating functional and integration tests
  • B2 level or higher English proficiency

Why the Work Is Challenging

Modern frontier AI models are already highly capable at programming. Designing tasks that can meaningfully challenge advanced coding models therefore requires a strong understanding of software development and the situations where AI systems tend to struggle.

The goal is to create scenarios that clearly distinguish strong solutions from weak ones. Because many tasks can have multiple valid implementations, designing tests that accept every correct approach while still rejecting incorrect solutions requires careful judgment and technical expertise.

How the Process Works

Apply → Complete qualification(s) → Join a project → Work on tasks → Receive payment

Compensation

Compensation can reach the equivalent of $50 per hour, depending on your experience level and working pace.

Individual tasks are estimated to require approximately 20 hours, and you have the flexibility to determine your own working schedule.

To apply for this job please visit jobs.workable.com.