Control Under Compression: Reliability Frontiers for Tool-Using Agents
Unknown
This paper introduces a framework for reliable tool-using agents under communication constraints, focusing on control compression and reliability frontiers.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces a framework for reliable tool-using agents under communication constraints, focusing on control compression and reliability frontiers.
Unknown
This paper introduces CAGE, a framework for certified authorization of tool-using LLM agents under typed-return uncertainty, ensuring safe runtime permission decisions.
Hankyul Baek, Jae-Koo Noh, Sanghyun Seo, et al.
This paper evaluates data leakage risks in tool-using LLM agents under realistic scenarios, proposing a framework to assess and mitigate such risks.
Unknown
MANTRA combines automated benchmark generation with SMT-based formal validation to enable scalable, reliable benchmarking of tool-using LLM agents.
Zhenting Wang, Qi Chang, Hemani Patel, et al.
Mcp-bench introduces a benchmark for evaluating tool-using LLM agents on complex real-world tasks via MCP servers, providing a standardized framework for assessing agent capabilities.
Dawei Li, Yuguang Yao, Zhen Tan, et al.
Introduces ToolPRMBench, a benchmark for evaluating process reward models (PRMs) in tool-using agents, and advances PRM training to improve agent performance.
Unknown
This paper investigates reinforcement learning for interactive tool-using agents, highlighting user model fine-tuning as a critical component for training such agents.
Unknown
MemToolAgent introduces a memory-augmented framework for tool-using agents that leverages environment and user feedback to improve general-purpose tool use and personalization.
Zhiqiang Liu, Wenhui Dong, Yilan Tan, et al.
TOBench is a task-oriented omni-modal benchmark for evaluating real-world tool-using agents through closed-loop multimodal verification.
Unknown
This paper presents a budgeted evaluation of ReAct variants for tool-using LLM agents, analyzing how compute constraints affect the Pareto frontier of accuracy versus inference cost.
Binjie Zhang, M. Shou
ReGRPO enhances tool-using agents by integrating reflection-guided correction into policy optimization, improving performance through iterative self-correction.
Xu Li, Simon Yu, Minzhou Pan, et al.
Introduces MTAgentRisk, the first multi-turn safety benchmark for tool-using agents, revealing a 16% average increase in attack success rate across multiple turns.