Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v1 Announce Type: cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2608. 15798v1 Announce Type: new Abstract: Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to.
arXiv:2607. 08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution.
arXiv:2606. 09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate.
arXiv:2608. 14588v1 Announce Type: new Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences.
arXiv:2603. 09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against.
arXiv:2605. 29055v2 Announce Type: replace Abstract: This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired Nested Learning architecture with Continuum Memory Systems (CMS) and semantic similarity caching, tested on a hybrid benchmark of 310 prompts (217 epistemic-uncertainty prompts, 93 fabrication-induction stress tests).
arXiv:2601. 02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctuates significantly depending on how information is distributed in real-world corpora.
arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.
arXiv:2608. 11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways.
arXiv:2608. 13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies.