arXiv:2609.34187v2 Announce Type: replace-cross
Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
By Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
The paper proposes measuring a language model’s understanding via no‑arbitrage, defining it as the inability of a bounded trader to profit from Dutch books against the model’s probabilities on logically related claims. It shows that full logical coherence is computationally infeasible, that standard next‑token training yields incoherent predictions across formats, and that uncertainty grows predictably along reasoning chains, creating arbitrage opportunities. The authors introduce Arbitr, a training framework that penalizes logical inconsistencies while maintaining accuracy, reducing exploitability by orders of magnitude and revealing a scaling illusion where large models appear coherent yet exhibit extreme unjustified confidence.
By Daniel Dragonevskiy
The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.
The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.
By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
arXiv:2609.37497v1 Announce Type: new
Abstract: Modern transformer models excel at capturing semantic relationships through sentence embeddings, yet their ability to perform pragmatic reasoning remai...
By Stefania Butnaru, Claudiu Creanga
arXiv:2601.19435v2 Announce Type: replace-cross
Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on s...
By Shengwei Xu, Zhaohua Chen, Xiaotie Deng, Zhiyi Huang, Grant Schoenebeck
The paper derives an information‑theoretic bound for a shared‑state cognitive architecture that uses an auxiliary variable to mediate context. It shows that the residual dependence of observable behavior on context, given the shared state, is bounded by the information carried by the auxiliary variable and its conditional entropy. A recognition‑memory example illustrates how to compute and compare this bound across different representational choices, providing a framework for analyzing context‑memory‑control trade‑offs in cognitive models and artificial agents.
By Song-Ju Kim
The paper critiques a recent NLI benchmark that tests the imperfective paradox, arguing that the benchmark suffers from conceptual and evaluation mis-specifications, notably Aspectual Reduction and a lack of strict NLI standards. The authors re-evaluate the benchmark, identify mis-specifications, and construct lexically matched minimal pairs to control for lexical variation. Their experiments reveal that models often exhibit a Sufficiency Bias, accept simple‑past hypotheses without affirming culmination, and that prompting interventions shift label decisions without improving true semantic understanding, highlighting additional failure modes such as compositional aspectual classification errors and surface‑form attraction.
By Kaiqiao Han, Yizhou Sun
arXiv:2605.27268v2 Announce Type: replace-cross
Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabu...
By Samer Awad, Javier Conde, Carlos Arriaga, Tairan Fu, Javier Coronado-Bl\'azquez, Pedro Reviriego
arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.
By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty...