The open benchmark for AI agent task execution. Claude Code vs Gemini CLI — who wins? Live leaderboard inside.
Real tasks. Real sandboxes. Real scores. No vibes.
Live Leaderboard | Methodology | Contributing
</div>Most agent benchmarks test toy problems or let agents self-report. AgentBench-Live drops agents into Docker-sandboxed workspaces with real codebases, real data files, and real multi-step workflows — then scores them automatically with test suites and LLM judges.
Why not SWE-bench / OpenHarness? Those benchmark a single axis (GitHub issue resolution). We test 5 capability domains — code, data analysis, multi-step orchestration, research, and tool use — because real work isn't just fixing bugs.
10 tasks across 5 domains | Docker sandbox | Auto-eval + LLM Judge scoring
| Agent | Code | Data | Multi-Step | Research | Tool Use | Overall |
|---|---|---|---|---|---|---|
| Claude Code | 1.00 | 0.07 | 0.74 | 0.60 | 0.25 | 0.53 |
| Gemini CLI | 1.00 | 0.32 | 0.77 | 0.45 | 0.05 | 0.52 |
| Codex CLI | - | - | - | - | - | pending |
| Aider | - | - | - | - | - | pending |
Universal, model-agnostic operating harness for AI agents (Claude, Codex, Gemini, …) — a lean core + work-type profiles assembled by one setup script.
Game-development Agent Skills for AI coding agents: install once and a master router loads the right skill for your engine and task. 66 original, version-pinned skills (plus a master router) in the portable SKILL.md format that runs across Claude Code, Cursor, Codex, Copilot, Gemini CLI and more, for Godot, Unity, Unreal, web and beyond.
A desktop pet for macOS & Windows that monitors your AI coding agents (Claude Code, Codex, Cursor, Gemini...) in real time, and grows as you code, feed it tokens, level it up, climb the leaderboard.
UltraGameStudio - AI coding agent for game development: engine workflows, gameplay code, and asset generation.
The coding agent that answers to you, your model, your machine, your rules.
Stop babysitting local AI agents. Just notifications, approve, and resume your Codex,Pi,Grok, or Claude code sessions anywhere. 0-Intrusion mobile control bridge via Telegram/微信/飞书. No hooks, no skills, no MCP.
Workflows from the Neura Market marketplace related to this Gemini resource