
AI Tools
VAKRA Benchmark Examines AI Agent Reasoning Failures
IBM Research details VAKRA, a benchmark that tests AI agents on reasoning and tool use in enterprise settings. It features over 8,000 APIs across 62 domains and four capabilities focused on API chaining, tool selection, multi-hop reasoning, and multi-source tasks with policies. Analysis shows models like GPT-OSS-120B lead but all struggle with complex workflows, as revealed in error breakdowns and performance charts.
Apr 155 minNeura News