Threats in LLM-Powered AI Agents Workflows
M. Ferrag, N. Tihanyi, Djallel Hamouda, et al.
A unified threat model for LLM-agent ecosystems covering host-to-tool and agent-to-agent attacks, with over 30 techniques and mitigation strategies.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
M. Ferrag, N. Tihanyi, Djallel Hamouda, et al.
A unified threat model for LLM-agent ecosystems covering host-to-tool and agent-to-agent attacks, with over 30 techniques and mitigation strategies.
K.A. Cobbina, Tianyi Zhou
First systematic study of how demo positions in prompts cause accuracy and prediction drift in LLMs, finding start-of-prompt placement yields +6 point gains.
Haozhe Wang, Qixin Xu, Che Liu, et al.
This paper identifies a two-phase reasoning hierarchy in LLMs trained with RL and proposes HICRA, a credit assignment algorithm that boosts performance by focusing optimization on high-level planning tokens.
S. Motwani, Alesia Ivanova, Ziyang Cai, et al.
Introduces a scalable method using curriculum RL on synthetically composed short-horizon data to boost long-horizon reasoning, achieving up to 2.06x accuracy gains on competition-level benchmarks.
Shuo Xing, Junyuan Hong, Yifan Wang, et al.
Shows continual pre-training on junk web text causes lasting cognitive decline in LLMs, with dose-response effects and partial healing.
Ahmed Heakl, Martin Gubri, Salman Khan, et al.
Dr. LLM retrofits frozen LLMs with lightweight routers trained via MCTS to skip, execute, or repeat layers, improving accuracy and efficiency without altering base weights.
James Y. Huang, Sheng Zhang, Qianchu Liu, et al.
Proposes BeMyEyes, a multi-agent framework that uses a small VLM as a perceiver and a text-only LLM as a reasoner to achieve multimodal reasoning without training large-scale models.
Jeremy Yang, Noah Yonack, Kathryn Zyskowski, et al.
First large-scale field study of general-purpose AI agent adoption, usage intensity, and use cases in open-world web environments.
Shyam Agarwal, Hao He, Bogdan Vasilescu
A longitudinal causal study finds that autonomous coding agents boost velocity only when adopted first, but persistently degrade software quality.
Unknown
Orca 2 introduces Cautious Reasoning to train smaller models to select effective solution strategies via task-specific system instructions.
Sarah Mohammed Alshehri, Sanaa Abdullah Sharaf, Rania Abdullrahman Molla
A systematic literature review of 28 studies (2020-2025) analyzing how graph neural networks are applied to detect cyberattacks in IoT, web, phishing, and network traffic domains.
P. Colombo, T. Pires, Malik Boudiaf, et al.
SaulLM-7B is a 7-billion-parameter LLM tailored for the legal domain, trained on over 30 billion tokens of English legal text and instruction-tuned to achieve state-of-the-art legal comprehension and generation.