arXiv Machine Learning By Marios Papamichalis, Regina Ruane

What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content

Read the original on arXiv Machine Learning →

The paper investigates how entropy over chain‑of‑thought tokens influences policy decisions such as gradient application, pruning, and collapse detection. By separating scaffold tokens from substantive content, the authors analyze entropy, Kullback–Leibler divergence, and entropy velocity for each channel, proving differences between raw and content conventions and bounding answer diversity. Empirical results across 23 configurations show that scaffold tokens can account for up to 41% of high‑entropy positions, with entropy share growing through distillation, while content conventions outperform raw surprisal on compression tasks and reveal significant answer leakage in re‑fed chains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv AI
Jun 19

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning

arXiv:2606. 19771v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a fundamental optimization instability: uniform token updates precipitate entropy collapse, leading to premature convergence to suboptimal strategies, whereas excessive Shannon Entropy maximization can cause entropy explosion, driving blind exploration toward incoherent reasoning chains.

By Xuanzhi Feng, Zhengyang Li, Zeyu Liu, Haoxi Li, Yuming Jiang, Bing Guo, Jingcai Guo, Jie Zhang, Song Guo