arXiv AI

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge

arXiv:2602. 02470v2 Announce Type: replace Abstract: Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" -- when trained on forward knowledge data of the form "$A \rightarrow B$" (e.

arXiv AI
Jul 28

LEDOM: Reverse Language Model

arXiv:2507. 01335v4 Announce Type: replace-cross Abstract: Autoregressive language models are trained exclusively left-to-right.

By Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan
arXiv AI
Aug 7

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

arXiv:2608. 06243v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level.

By ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng
arXiv Computation and Language
Sep 11

Probing for Knowledge Attribution in Large Language Models

The paper introduces a method for identifying the dominant knowledge source behind large language model (LLM) outputs, distinguishing between faithfulness violations (misuse of provided context) and factuality violations (errors in internal knowledge). A simple linear probe trained on hidden representations can reliably classify this source, and the authors present AttriWiki, a self‑supervised pipeline that generates labeled training data by prompting models to recall withheld entities or read them from context. Probes trained on AttriWiki achieve high Macro‑F1 scores across several models and datasets, generalize zero‑shot to a benchmark, and show that attribution mismatches can increase error rates by up to 70%. "whyItMatters":"The study demonstrates that knowing the source of an LLM’s answer is crucial for effective mitigation of hallucinations, as attribution mismatches significantly raise error rates."

By Ivo Brink, Alexander Boer, Dennis Ulmer
arXiv Computation and Language
Sep 18

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery is a self‑supervised method that turns failed reasoning attempts into training data, enabling large language models to learn how to correct mistakes during inference. By extracting initial segments of erroneous trajectories and using them as prompts, the approach teaches models to recognize and recover from errors without external critics. Experiments show significant accuracy gains on benchmarks such as AIME 2025 and Minerva, and the method overcomes the scaling collapse problem, fostering emergent self‑correction behaviors.

By Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong
arXiv Machine Learning
Sep 11

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Negative Self-Distillation (NSD) is a new framework for improving large language models by encouraging them to diverge from their own flawed reasoning rather than imitate privileged solutions. Unlike On-Policy Self-Distillation, which can suppress uncertainty and exploratory behavior, NSD generates a question‑specific negative condition (e.g., a careless reasoner) and uses a dynamic gating mechanism to target only reasoning‑critical tokens for penalization. This approach preserves foundational language capabilities while consistently outperforming OPSD and other label‑free self‑bootstrapping reinforcement learning baselines.

By Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng
arXiv Computation and Language
Aug 28

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

Large language models can solve complex multi‑hop tasks but often fail on simple two‑hop queries, even when each hop is individually correct. In a controlled symbolic setting, the authors find that models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Mechanistic analysis shows that successful generalization relies on consistent intermediate representations across contexts, whereas failures arise from a mismatch between lower‑layer representation construction and upper‑layer mapping to outputs. The study proposes a recurrent‑style training strategy that improves out‑of‑distribution two‑hop generalization.

By Zili Zhang, Yilin Wang, Heng Wang, Herun Wan, Minnan Luo
arXiv Computation and Language
Sep 17

Multi-Hop Knowledge Composition is Bound by Pretraining Exposure

Large Language Models struggle with implicit multi‑hop reasoning, correctly answering individual facts but failing to combine them in a single pass. In a controlled setting, the authors show that this failure persists even with high 1‑hop accuracy, indicating it is due to pretraining exposure rather than missing knowledge. They test nine data‑centric augmentation formats and find that only individuals seen in compositional contexts during pretraining enable transfer to unseen questions, proving exposure to such contexts is necessary for implicit multi‑hop reasoning.

By Yannis Karmim, Luis Marti, Djam\'e Seddah, Valentin Barri\`ere