arXiv Machine Learning By Zhenyuan Guo, Tong Chen, Wenlong Meng, Chen Gong, Xin Yu, Chengkun Wei, Wenzhi Chen

Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

Read the original on arXiv Machine Learning →

arXiv:2601. 18383v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 15

Fractured Chain-of-Thought Reasoning

arXiv:2505. 12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining.

By Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong
arXiv Machine Learning
Sep 7

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

BeaconKV is a training‑free key‑value cache compression technique for Large Reasoning Models that uses beacon queries—compact representatives of query clusters—to predict which KV pairs will be revisited during long‑horizon reasoning. By focusing on Thought Revisiting Tokens that re‑attend distant context, BeaconKV reduces memory usage up to 5.8× and improves throughput by over 4.3× while largely preserving cache accuracy across multiple open‑source LRMs and reasoning benchmarks.

By Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi
arXiv Computation and Language
Sep 10

Demystifying Entropy-based Selection for Chain-of-Thought Compression in Large Reasoning Models

The paper evaluates entropy-based pruning for compressing Chain-of-Thought (CoT) reasoning in large models. Across multiple models and tasks, low- and high-entropy step selection shows no advantage over random pruning, and low-entropy token retention only helps on mathematical benchmarks due to the low entropy of numeric tokens. Patching a few CoT tokens with their original activations restores near-perfect performance, indicating that task information is distributed throughout the entire reasoning chain rather than concentrated in a small set of tokens.

By Sara Candussio, Daniel Scalena, Luca Bortolussi, Elisabetta Fersini, Malvina Nissim, Gabriele Sarti