arXiv:2609.37497v1 Announce Type: new
Abstract: Modern transformer models excel at capturing semantic relationships through sentence embeddings, yet their ability to perform pragmatic reasoning remai...
By Stefania Butnaru, Claudiu Creanga
arXiv:2603. 01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning.
By Kyle Cox, Darius Kianersi, Adri\`a Garriga-Alonso
arXiv:2603.29396v2 Announce Type: replace
Abstract: Standard evaluations of Large language models (LLMs) focus on task performance, offering limited insight into whether correct behavior reflects app...
By Zo\"e Prins, Samuele Punzo, Frank Wildenburg, Giovanni Cin\`a, Sandro Pezzelle
arXiv:2606. 15733v1 Announce Type: cross Abstract: Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer are unchanged.
By Zhenyu Yu
arXiv:2609.39225v1 Announce Type: new
Abstract: Argument structure prediction (ASP) constructs complete argument structures from discourse by identifying argumentative units and their relations. Whil...
By Siddharth Bhargava, Sara Tonelli, Patricia Mart\'in-Rodilla, Javier Parapar
The paper investigates how adjectival modifiers affect the semantic plausibility of events, using the Adept benchmark of 16,000 English sentence pairs that differ by a single adjective. Experiments show that sentence transformers, despite being conceptually suited to the task, underperform compared to models like RoBERTa. The authors provide an error analysis and discuss the implications of their findings for future work on balancing training and test data.
By Anna Golub, Beate Zywietz, Annerose Eichel
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both.
arXiv:2601. 16407v3 Announce Type: replace-cross Abstract: Large language models (LLMs) make next-token predictions based on clues present in their context, such as semantic descriptions and in-context examples.
By Toni J. B. Liu, Baran Zadeo\u{g}lu, Nicolas Boull\'e, Rapha\"el Sarfati, Gurbir Arora, Christopher J. Earls
arXiv:2604. 22128v2 Announce Type: replace-cross Abstract: When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering.
By Aryan Sharma, Cutter Dawes, Shivam Raval
arXiv:2607. 21498v1 Announce Type: cross Abstract: A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen {\guillemotleft}This is not a course.
By Federico Boggia
arXiv:2608.30413v1 Announce Type: new
Abstract: Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of n...
By Jayanta Sadhu, Sayem Shahad, Kenneth Marino
arXiv:2504. 09762v4 Announce Type: replace Abstract: Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks.
By Subbarao Kambhampati, Karthik Valmeekam, Siddhant Bhambri, Vardhan Palod, Lucas Saldyt, Kaya Stechly, Soumya Rani Samineni, Durgesh Kalwar, Upasana Biswas