Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction
Read the original on arXiv Machine Learning →This study independently reproduces the dissociation reported by Zhao (2026) regarding chain-of-thought entropy in large language models. It confirms that the shape of the entropy trajectory predicts answer correctness, while the total entropy drop magnitude does not, across four open-weight models and two benchmarks (GSM8K and MATH‑500). The reproduction also maps settings where the magnitude signal holds or fails and documents protocol differences not reported in the original work.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.