Control Under Compression: Reliability Frontiers for Tool-Using Agents
Unknown
This paper introduces a framework for reliable tool-using agents under communication constraints, focusing on control compression and reliability frontiers.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces a framework for reliable tool-using agents under communication constraints, focusing on control compression and reliability frontiers.
Neura Market
81% of n8n workflow templates that call an external system declare no error handling at all — no retry, no timeout, no error branch. A structural analysis of 5,147 deduplicated automation templates, read from their own executable specifications.
Fali Wang, Zhiwei Zhang, Xianren Zhang, et al.
This survey systematically defines Small Language Models (SLMs), provides a taxonomy of methods for their acquisition, application, enhancement, and reliability, and offers frameworks for their effective use in resource-constrained settings.
Jiawei Gu, Xuhui Jiang, Zhichao Shi, et al.
This survey systematically addresses how to build reliable LLM-as-a-Judge systems, covering strategies to enhance reliability, evaluation methodologies, and a novel benchmark.
J. Jakeman, Lorena A. Barba, Joaquim R. R. A. Martins, et al.
This paper proposes a framework for verification and validation (V&V) of scientific machine learning models to enhance trustworthiness and reliability in scientific applications.
Anand Shankar, Bikash Chandra Sahana
This paper proposes an ensemble of machine learning approaches (boosting, bagging, stacking) to predict low visibility and dense fog at Patna airport, achieving high reliability for aviation services.
Maxime M'eloux, Silviu Maniu, Franccois Portet, et al.
This paper investigates whether mechanistic interpretability methods are identifiable, proposing a framework to assess the uniqueness and reliability of explanations derived from neural network internals.
Unknown
This paper outlines a research paradigm for scalable oversight of large language models, based on Cotra's sandwiching approach, to improve model reliability.
Unknown
This paper proposes a robust method for detecting hallucinations in large language model outputs, improving reliability of question answering systems.
Unknown
This paper identifies distribution shift and scale as failure modes of benchmark contamination detection, revealing a reliability gap in benchmark auditing.
Ana Reyna, Cristian Martín, Jaime Chen, et al.
This paper surveys the integration of blockchain with IoT, analyzing challenges and opportunities for enhancing security and data reliability in distributed IoT environments.
Peilin Feng, Suorong Yang, Soujanya Poria
Introduces Σ-Mem, an online reliability memory for LLM-based multi-agent systems that records and updates agent competence and relationship evidence, enabling stable adaptation and improved coordination.