arXiv AI By Kushal Chakrabarti

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models

Read the original on arXiv AI →

arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction

This study independently reproduces the dissociation reported by Zhao (2026) regarding chain-of-thought entropy in large language models. It confirms that the shape of the entropy trajectory predicts answer correctness, while the total entropy drop magnitude does not, across four open-weight models and two benchmarks (GSM8K and MATH‑500). The reproduction also maps settings where the magnitude signal holds or fails and documents protocol differences not reported in the original work.

By Theodore O. Cochran