Freelance Agent Evaluation Engineer at Mindrift — AI Jobs | Neura Market
    Neura Market
    Neura Market
    /Jobs
    Marketplace
    Directories
    Resources
    AI JobsFreelance Agent Evaluation Engineer
    Mindrift

    Freelance Agent Evaluation Engineer

    Mindrift

    Australia

    Marketplace

    • Prompts
    • Workflows
    • Agent Hub
    • Workflow Packs
    • Categories
    • Marketplace

    Directories

    • AI Tools Directory
    • ChatGPT
    • Claude
    • Gemini
    • Cursor
    • Grok
    • DeepSeek
    • Perplexity
    • CoPilot
    • Midjourney
    • Stable Diffusion
    • MCP Servers
    • .md Directory
    • All Directories

    Free Tools

    • AI Text Humanizer
    • AI Content Detector
    • Workflow Generator
    • Model Comparison
    • AI Pricing Calculator
    • AI Benchmarks
    • ROI Calculator
    • All Free Tools

    Resources

    • AI News
    • Blog
    • AI Answers
    • Error Solutions
    • AI Tutorials
    • AI Agent Guides
    • AI Models
    • AI Research Papers
    • Integrations
    • Alternatives
    • n8n vs Zapier
    • Make vs Zapier
    • n8n vs Make
    • Resource Library
    • Documentation
    • API Access to Our Data

    Community

    • AI Newsletter
    • AI Jobs
    • AI Events
    • AI Companies
    • Start Selling
    • Sell n8n Workflows
    • Sell AI Agents
    • Sell Prompts
    • Creator Guide
    • Advertise
    • Affiliates

    Company

    • About
    • Contact
    • Help
    • Careers
    • Pricing
    • Terms
    • Privacy
    • License
    • DMCA

    The #1 Newsletter in AI

    Weekly updates, news, and content that matter.

    Neura Market Logoneuramarket

    © 2026 Neura Market. All rights reserved.

    AUD 100,000
    Mid-level / Intermediate
    Full-time
    Remote
    8/3/2026
    Apply

    About This Role

    Please submit your CV in English and indicate your level of English proficiency.

    Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

    What this opportunity involves 

    We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

    You'll create challenging tasks and evaluation criteria within realistic simulated environments:

    • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
    • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
    • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
    • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

    What this is NOT

    • Not data labeling
    • Not prompt engineering
    • Not writing code from scratch - the agent writes most of the code; you guide and evaluate

    What we look for

    • 5+ years in software development
    • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
    • Experience writing tests (functional, integration)
    • English proficiency - B2+

    Why this is hard 

    Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

    How it works

    Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

    Effort estimate

    Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

    Compensation

    Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.

    Skills & Tech Stack

    PythonTypeScriptJavaScriptDockerReactPostgreSQLRedisApache Kafka

    Roles

    Engineer

    Location

    Region

    Asia/Pacific

    Country

    Australia

    Topics

    QA & Testing

    Related AI Jobs

    Turnitin, LLC

    Senior Software Quality Engineer (Mexico Remote)

    Turnitin, LLC·Full-time·Mexico
    QA & Testing
    Dataiku

    Software Engineer in Test - Onsite or Remote (FR, UK, DE, NL)

    Dataiku·Full-time·Europe, France, Germany, Netherlands, UK
    QA & Testing
    Truelogic

    Senior SDET - GovTech Industry (Colombia)

    Truelogic·Full-time·LATAM
    QA & Testing
    Actian

    QA Automation Lead [gn] Data Intelligence

    Actian·Full-time·Europe
    QA & Testing
    ← Back to all jobs