500 Bankers Review AI Outputs, None Fit for Clients
A fresh benchmark challenged leading AI models such as GPT-5.4 and Claude Opus 4.6 with daily tasks faced by junior investment bankers. After review by 500 bankers, zero outputs qualified for direct client use. More than half the reviewers, however, indicated they would employ the results as an initial draft.
Researchers from Handshake AI and McGill University unveiled BankerToolBench, a public benchmark assessing AI agents on standard junior investment banker processes. Handshake AI operates as the commercial side of Handshake, a career site that assigns screened academics and experts to AI labs for model training and assessment. Nine prominent models underwent the evaluation, and the bankers delivered a clear judgment: no results suit client presentation.
Benchmark Design and Banker Input
Bankers reported that 41 percent of AI-generated content required substantial revisions, while 27 percent proved entirely useless. Only 13 percent needed minor adjustments, and none earned approval without changes.
The project involved roughly 500 active and past investment bankers from places like Goldman Sachs, JPMorgan, Evercore, Morgan Stanley, and Lazard. Among them, 172 crafted the tasks, contributing over 5,700 hours. The 100 tasks averaged five hours per human banker, with some extending to 21 hours.
BankerToolBench evaluates tangible products a junior banker submits to a manager: functional Excel financial models, PowerPoint presentations for clients, PDF reports, and Word documents.
Agents must search data rooms, access platforms such as FactSet and Capital IQ, and analyze SEC filings. One task might involve up to 539 language model calls, 97 percent linked to tools or code runs.
Reviewers apply a banker-created checklist with about 150 checks across six categories: technical accuracy, client suitability, compliance, verifiability, and file consistency.
An AI checker named Gandalf, built on Gemini 3 Flash Preview, handles scoring. It matches human evaluators 88.2 percent of the time, exceeding the 84.6 percent match between two humans.
Model Results and Shortfalls
Tests covered GPT-5.2, GPT-5.4, Claude Opus 4.5 and 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview, Grok 4, plus open models Qwen-3.5-397B and GLM-5. GPT-5.4 performed best yet missed nearly half the standards. Just 16 percent of its work served as a viable base; with three consistent trials required, that fell to 13 percent.
No model produced a single client-ready output. For GPT-5.4, 2 percent of tasks met all key weighted standards. Gemini 2.5 Pro achieved zero.
Claude Opus 4.6 generates sleek-looking results initially. Yet Excel files often embed key figures as static numbers instead of formulas, blocking scenario tests. Altering the buyout price yields no updates. Claude Opus 4.5 shared this issue.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
GPT-5.4 earned 58.1 out of 100 total and outperformed GPT-5.2 in 70 percent of direct comparisons. Claude Opus 4.6 and Gemini 3.1 Pro scored close behind, with Grok 4 and Gemini 2.5 Pro lagging.
Failure Patterns and Examples
GPT-5.4 agent paths revealed four main errors. Code and formula bugs hit 41 percent: agents invoke nonexistent python-pptx functions and delete faulty code rather than correct it.
Business logic failed in 27 percent, like placing cost synergies in revenue. Data fetches aborted in 18 percent. Agents invented 13 percent of absent data as authentic.
Claude Opus 4.6 topped Client Readiness at 63 points and Risk & Compliance at 46. It managed 47 on Technical Correctness, trailing GPT-5.4's 57.
Paper examples show sneaky mistakes. One deck listed revenue at $189.5 billion on one slide, $201.0 billion on another for the same timeframe.
Another picked Netflix red accents against a bank blue policy. In pharma deal review, an agent made up trial data absent from SEC records.
Models fared better on PowerPoint than Excel. Debt capital markets, merger models, and capital structure tables posed biggest hurdles. Added assumed context boosted scores.
Broader Context and Limits
BankerToolBench supports reinforcement learning too. Qwen-3-4B and 32B saw five- to thirteen-fold gains via Dr. GRPO and DPO, from low starts.
Drawbacks include US focus, no private deal data, and omission of bank team iterations. Still, creators view it as a thorough check on AI for complex knowledge jobs. All materials sit public.
Results echo others. Vals.ai with a major bank found OpenAI o3 at 48.3 percent on finance analysis. UC Berkeley noted production agents use basic, controlled steps. Carnegie Mellon and Stanford said agent work overlooks finance, law, management.
Anthropic addresses gaps: Claude now toggles Excel and PowerPoint independently. Cowork plugins feed FactSet, MSCI, LSEG data into flows.
Related on Neura Market

