Reinforcement Learning via Self-Play
Guanghao Ye, Khiem Pham, Xinzhi Zhang, et al.
Proposes RLSP, a post-training framework that decouples exploration and correctness signals during PPO to enable emergent reasoning behaviors in LLMs.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Guanghao Ye, Khiem Pham, Xinzhi Zhang, et al.
Proposes RLSP, a post-training framework that decouples exploration and correctness signals during PPO to enable emergent reasoning behaviors in LLMs.
Roeland Scheepens, Niels Willems, Huub van de Wetering, et al.
This paper introduces a flexible architecture for multivariate trajectory density maps using composable blocks and expressions, enabling analysts to model domain knowledge for visual exploration.
Bingxin Xu, Yuzhang Shang, Emilio Ferrara
BATON improves long-horizon robot manipulation by making subtasks the unit of exploration and adding transition-aware memory, boosting task success by 11.6% over SoTA.
Bogdan Georgiev, J. Gómez-Serrano, Terence Tao, et al.
This paper demonstrates AlphaEvolve, an LLM-guided evolutionary coding agent, as a tool for autonomously discovering novel mathematical constructions, rediscovering best-known solutions and improving several across 67 problems.
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
Introduces parameter-space exploration for LLM RL via Perturbed Parameter Policy Optimization (3PO), improving downstream performance over GRPO at similar FLOPs.
Unknown
This paper introduces a safe model-based reinforcement learning approach that guarantees stability during exploration and learning.
Unknown
This paper presents Gentron, a diffusion transformer architecture adapted from class to text conditioning for image and video generation, with empirical exploration of conditioning mechanisms.
Zhen Fang, Yu Zeng, Wenxuan Huang, et al.
Video-DeepResearch extends multimodal agents to continuous video streams with a decoupled perception-exploration pipeline, achieving SOTA 64.0% on a new benchmark.
Unknown
This paper provides a comprehensive exploration of LLM evaluation from a metrics perspective, offering a pragmatic guide for effective metric selection.
Unknown
This paper systematically categorizes and explores KV cache compression techniques for transformer-based models, providing a structured analysis of methods to reduce memory overhead.
Unknown
This paper explores in-context learning at extreme scale with long-context models, studying properties of ICL and long-context behavior across multiple datasets.
Unknown
ONELIFE learns symbolic world models from a single unguided exploration episode using programmatic representations and stochastic abstractions.