Lookahead Routing for LLMs
Canbin Huang, Tianyuan Shi, Yuhua Zhu, et al.
Lookahead routing predicts latent representations of potential model outputs to guide LLM selection, improving routing decisions without full inference.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Canbin Huang, Tianyuan Shi, Yuhua Zhu, et al.
Lookahead routing predicts latent representations of potential model outputs to guide LLM selection, improving routing decisions without full inference.
Chirag Vashist, Shichong Peng, Ke Li
RS-IMLE improves few-shot image synthesis by redesigning the prior distribution to align training and inference latent codes, yielding higher quality images than existing GAN and IMLE methods.
Daman Arora, Andrea Zanette
This paper proposes using reinforcement learning to train large reasoning models to dynamically allocate inference-time compute based on task complexity, reducing inference costs while preserving accuracy.
Xiaowei Cai, Yunuo Cai, Bingao Chen, et al.
τ_0-VLA is a hierarchical robot foundation model that uses world-model-guided test-time computation to allocate additional inference compute for high-level subtask generation, improving long-horizon manipulation.
Zeyu Ren, Ling Yue, Ran Li, et al.
FlowEvo is a training-free framework where workflows and executable skills co-evolve at inference time, improving agent performance across benchmarks while reducing token usage.
Rujikorn Charakorn, Edoardo Cetin, Shinnosuke Uesaka, et al.
Doc-to-LoRA (D2L) is a lightweight hypernetwork that meta-learns to perform approximate context distillation in a single forward pass, generating LoRA adapters for a target LLM to reduce inference latency and KV-cache memory.
Keming Wu, Baoyi Wang, Kaichen Zhang, et al.
StreamOPD is a post-training recipe combining verifiable streaming-video data, thinking-mode on-policy distillation, and instruct-mode deployment, improving streaming video understanding without inference-time memory or retrieval.
Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, et al.
Introduces Internalized Visual Thinking (IVT), a post-training framework that predicts latent future frames during training to enable efficient, direct-answer video reasoning at inference, outperforming Visual CoT with 5x lower latency.
Lijie Yang, Hongyin Luo, Jiawei Zhao, et al.
Gambit is an inference algorithm that performs thought-level beam search to dynamically allocate test-time compute to the most promising reasoning traces, improving accuracy and throughput while reducing token consumption.
Erik Johannes Husom, Arda Göknil, Merve Astekin, et al.
This paper evaluates 28 quantized LLMs on a Raspberry Pi 4, measuring energy efficiency, accuracy, and latency to identify optimal configurations for sustainable edge AI deployment.
Ebenezer Gelo, Geraud Nangue Tasse, Steven James, et al.
Proposes RCI framework to convert sparse trajectory-level stop-feedback into dense per-step costs for safe offline RL, preserving feasible policy set and optimal Lagrangian.
Cheng Qian, Wenting Zhao, Liangwei Yang, et al.
This paper introduces test-time strong-to-weak capability transfer via inference-time harnesses, nearly doubling target model performance on Theory-of-Mind benchmarks without parameter updates.