Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
This study independently reproduces the dissociation reported by Zhao (2026) regarding chain-of-thought entropy in large language models. It confirms that the shape of the entropy trajectory predicts answer correctness, while the total entropy drop magnitude does not, across four open-weight models and two benchmarks (GSM8K and MATH‑500). The reproduction also maps settings where the magnitude signal holds or fails and documents protocol differences not reported in the original work.
arXiv:2607. 08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution.
arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
arXiv:2604. 08941v2 Announce Type: replace Abstract: Medical Vision-Language Models (VLMs) answering binary presence questions on chest radiographs can fail in two linked ways: they are confidently wrong, and they change answers when a clinically equivalent question is rephrased.
arXiv:2609.07901v1 Announce Type: new Abstract: Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4...
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.