LangMARL: Natural Language Multi-Agent Reinforcement Learning (April 2026)
FreeBrings credit assignment and policy gradient evolution from cooperative MARL into language space — enables LLM agents to autonomously evolve coordination strategies in dynamic environments
About LangMARL: Natural Language Multi-Agent Reinforcement Learning (April 2026)
LangMARL is a framework that addresses the challenge of LLM-based agents autonomously evolving coordination strategies in dynamic environments. It identifies a multi-agent credit assignment bottleneck in LLM systems and brings credit assignment and policy gradient evolution from cooperative multi-agent reinforcement learning (MARL) into the language space. Key innovations include agent-level language credit assignment, gradient evolution in language space for policy improvement, and summarization of task-relevant causal relations from replayed trajectories to provide dense feedback and improve convergence under sparse rewards. Experiments across diverse cooperative multi-agent tasks demonstrate improved sample efficiency, interpretability, and strong generalization.
Key Features
Pros & Cons
- Addresses a key bottleneck in LLM-based multi-agent systems: credit assignment
- Demonstrates improved sample efficiency over baseline methods
- Offers interpretable policy improvements via language-space reasoning
- Generalizes well across diverse cooperative tasks