Freelance Agent Evaluation Engineer at Mindrift — AI Jobs | Neura Market
    Neura Market
    Neura Market
    /Jobs
    Marketplace
    Directories
    Resources
    AI JobsFreelance Agent Evaluation Engineer
    Mindrift

    Freelance Agent Evaluation Engineer

    Mindrift

    LATAM

    Marketplace

    • Prompts
    • Workflows
    • Agent Hub
    • Workflow Packs
    • Categories
    • Marketplace

    Directories

    • AI Tools Directory
    • ChatGPT
    • Claude
    • Gemini
    • Cursor
    • Grok
    • DeepSeek
    • Perplexity
    • CoPilot
    • Midjourney
    • Stable Diffusion
    • MCP Servers
    • .md Directory
    • All Directories

    Free Tools

    • AI Text Humanizer
    • AI Content Detector
    • Workflow Generator
    • Model Comparison
    • AI Pricing Calculator
    • AI Benchmarks
    • ROI Calculator
    • All Free Tools

    Resources

    • AI News
    • Blog
    • AI Answers
    • Error Solutions
    • AI Tutorials
    • AI Agent Guides
    • AI Models
    • AI Research Papers
    • Integrations
    • Alternatives
    • n8n vs Zapier
    • Make vs Zapier
    • n8n vs Make
    • Resource Library
    • Documentation
    • API Access to Our Data

    Community

    • AI Newsletter
    • AI Jobs
    • AI Events
    • AI Companies
    • Start Selling
    • Sell n8n Workflows
    • Sell AI Agents
    • Sell Prompts
    • Creator Guide
    • Advertise
    • Affiliates

    Company

    • About
    • Contact
    • Help
    • Careers
    • Pricing
    • Terms
    • Privacy
    • License
    • DMCA

    The #1 Newsletter in AI

    Weekly updates, news, and content that matter.

    Neura Market Logoneuramarket

    © 2026 Neura Market. All rights reserved.

    USD 80,000
    Mid-level / Intermediate
    Full-time
    Remote
    7/13/2026
    Apply

    About This Role

    Please submit your CV in English and indicate your level of English proficiency.

    Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

    What this opportunity involves 

    We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

    You'll create challenging tasks and evaluation criteria within realistic simulated environments:

    • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
    • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
    • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
    • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

    What this is NOT

    • Not data labeling
    • Not prompt engineering
    • Not writing code from scratch - the agent writes most of the code; you guide and evaluate

    What we look for

    • 5+ years in software development
    • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
    • Experience writing tests (functional, integration)
    • English proficiency - B2+

    Why this is hard 

    Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

    How it works

    Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

    Effort estimate

    Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

    Compensation

    Up to $40/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.

    Skills & Tech Stack

    PythonTypeScriptJavaScriptDockerReactPostgreSQLRedisApache Kafka

    Roles

    Engineer

    Location

    Country

    LATAM

    Topics

    Software Engineering

    Related AI Jobs

    dscout

    AI Data Engineer

    dscout·Full-time·USA
    Software Engineering
    Meta

    Software Engineer (Technical Leadership)

    Meta·Full-time·USA
    Software Engineering
    Cresta

    Senior Software Engineer, Backend - Platform Team

    Cresta·Full-time·Canada
    Software Engineering
    Zartis

    Senior AI Software Engineer

    Zartis·Full-time·Europe
    Software Engineering
    Cloudbeds

    Staff Software Engineer

    Cloudbeds·Full-time·Europe
    Software Engineering
    saas.group

    Head of Engineering

    saas.group·Full-time·Europe
    Software Engineering
    ← Back to all jobs