Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v1 Announce Type: cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
arXiv:2608. 20290v1 Announce Type: new Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses.
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts.
arXiv:2610.00694v1 Announce Type: cross Abstract: Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on t...
arXiv:2608. 13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies.
arXiv:2607. 14112v1 Announce Type: cross Abstract: Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale.
arXiv:2606. 09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate.
arXiv:2608. 15798v1 Announce Type: new Abstract: Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to.
The paper argues that large language models cannot achieve perfect reliability for any task, even with unlimited scale. It establishes that each generative task has an inherent reliability ceiling set by how much output uncertainty can be resolved from observable context, with a resolvable part that can be improved by more context and a subjective part tied to task ambiguity. The authors derive a scaling law showing that performance is limited by the scarcer resource—either training data or model capacity—and explain how this law explains phenomena such as retrieval augmentation and catastrophic forgetting.
The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.
The paper audits a developer‑accessible on‑device language model, revealing that it can confidently produce incorrect answers while refusing benign prompts, a phenomenon termed task‑asymmetric miscalibration. The model’s confident outputs are surface‑indistinguishable, with classifiers based on user‑visible features failing to separate correct from wrong responses. The authors propose a model‑agnostic audit protocol, a surface‑indistinguishability test, and a black‑box consistency wrapper that improves reliability without requiring model access.