arXiv Machine Learning

Lasting Effects of Abstract Pretraining Beyond Perplexity

arXiv Computation and Language
Sep 17

Multi-Hop Knowledge Composition is Bound by Pretraining Exposure

Large Language Models struggle with implicit multi‑hop reasoning, correctly answering individual facts but failing to combine them in a single pass. In a controlled setting, the authors show that this failure persists even with high 1‑hop accuracy, indicating it is due to pretraining exposure rather than missing knowledge. They test nine data‑centric augmentation formats and find that only individuals seen in compositional contexts during pretraining enable transfer to unseen questions, proving exposure to such contexts is necessary for implicit multi‑hop reasoning.

By Yannis Karmim, Luis Marti, Djam\'e Seddah, Valentin Barri\`ere
arXiv Machine Learning
4d ago

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...

By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko
arXiv Computation and Language
Aug 28

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

Large language models can solve complex multi‑hop tasks but often fail on simple two‑hop queries, even when each hop is individually correct. In a controlled symbolic setting, the authors find that models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Mechanistic analysis shows that successful generalization relies on consistent intermediate representations across contexts, whereas failures arise from a mismatch between lower‑layer representation construction and upper‑layer mapping to outputs. The study proposes a recurrent‑style training strategy that improves out‑of‑distribution two‑hop generalization.

By Zili Zhang, Yilin Wang, Heng Wang, Herun Wan, Minnan Luo
arXiv Computation and Language
Sep 22

SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine

arXiv:2410.17021v2 Announce Type: replace Abstract: Large Language Models with chain-of-thought prompting, such as OpenAI-o1, have shown impressive capabilities in natural language inference tasks. H...

By Xiaochen Wang, Liang Chen, Reza Haf Zhe Yang, Yiru Wang, Xiangdi Meng, Kunhao Pan, Zhifang Sui, Junqing He