PreprintarXiv.org2025
Fine-tuning LLM Agents without Fine-tuning LLMs
Huichi Zhou, Yihang Chen, Siyuan Guo, et al.
Introduces a memory-based online reinforcement learning paradigm for LLM agents that enables continual adaptation without fine-tuning the underlying LLM.
81Aug 22, 2025Large Language ModelsFine Tuning
arXiv