BeaconKV is a training‑free key‑value cache compression technique for Large Reasoning Models that uses beacon queries—compact representatives of query clusters—to predict which KV pairs will be revisited during long‑horizon reasoning. By focusing on Thought Revisiting Tokens that re‑attend distant context, BeaconKV reduces memory usage up to 5.8× and improves throughput by over 4.3× while largely preserving cache accuracy across multiple open‑source LRMs and reasoning benchmarks.
By Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi
arXiv:2505. 12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining.
By Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong
The paper introduces Planned Test-Time Scaling (PTTS), a method that replaces independent sampling of reasoning branches with a coordinated joint policy. PTTS uses a planner to generate distinct solution outlines for each branch and an executor to produce full solutions, thereby improving coverage of complementary reasoning modes. Two variants—PTTS‑ZS (zero‑shot) and PTTS‑RL (reinforcement‑learned)—demonstrate significant gains on five mathematical reasoning benchmarks, with PTTS‑RL achieving up to a 13.4‑point improvement in pass@64 over repeated sampling.
By Xueqing Wu, Langxing Bai, Hritik Bansal, Po-Nien Kung, Shuo Li, Hao Liu, Nanyun Peng, Kai-Wei Chang
arXiv:2608.21614v1 Announce Type: new
Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases de...
By Yujie Zhang, Bin Gao, Tulika Mitra
arXiv:2609.39334v1 Announce Type: cross
Abstract: Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, subs...
By Jinwoo Jeong (Korea University), Woohyung Choi (Korea University), Myeongjae Jeon (POSTECH), Jeongseob Ahn (Korea University)
Parason is a new framework that discovers and exploits both subtask and trial parallelism in large language model (LLM) reasoning. By converting sequential reasoning traces into structured parallel trajectories and training with Parallelism-Aware Group Relative Policy Optimization, it balances accuracy, latency, and parallelism. Experiments on mathematical reasoning benchmarks such as AIME24 and AIME25 show that Parason achieves an average acceleration of about 1.7× while maintaining competitive accuracy.
By Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han, Ligeng Zhu