Multi-Hop Knowledge Composition is Bound by Pretraining Exposure
Large Language Models fail at implicit multi-hop reasoning: a model answers "When was $X$ born? " and "Who is $Y$'s closest friend?
Large Language Models struggle with implicit multi‑hop reasoning, correctly answering individual facts but failing to combine them in a single pass. In a controlled setting, the authors show that this failure persists even with high 1‑hop accuracy, indicating it is due to pretraining exposure rather than missing knowledge. They test nine data‑centric augmentation formats and find that only individuals seen in compositional contexts during pretraining enable transfer to unseen questions, proving exposure to such contexts is necessary for implicit multi‑hop reasoning.
Large Language Models fail at implicit multi-hop reasoning: a model answers "When was $X$ born? " and "Who is $Y$'s closest friend?
arXiv:2609.38764v1 Announce Type: new Abstract: Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract,...
Large language models can solve complex multi‑hop tasks but often fail on simple two‑hop queries, even when each hop is individually correct. In a controlled symbolic setting, the authors find that models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Mechanistic analysis shows that successful generalization relies on consistent intermediate representations across contexts, whereas failures arise from a mismatch between lower‑layer representation construction and upper‑layer mapping to outputs. The study proposes a recurrent‑style training strategy that improves out‑of‑distribution two‑hop generalization.
arXiv:2607. 00341v1 Announce Type: cross Abstract: Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT).
arXiv:2604. 07822v2 Announce Type: replace-cross Abstract: We study implicit reasoning, i.
The paper investigates how large language models (LLMs) perform multi‑hop reasoning and challenges the prevailing hop‑aligned circuit hypothesis, which posits that bridge entities are computed sequentially across layers. Through systematic analyses of real‑world multi‑hop queries, the authors discover a phenomenon called layer‑order inversion, where later‑hop answer entities become decodable earlier than bridge entities, and this effect grows with the number of hops. They propose a probabilistic recall‑and‑extract framework that models multi‑hop reasoning as broad probabilistic recall in shallow MLP layers followed by selective extraction in deeper attention layers, and validate this framework with probing analyses that reinterpret prior evidence, explain chain‑of‑thought gains, and diagnose multi‑hop failures.
arXiv:2504. 03635v4 Announce Type: replace Abstract: Reasoning is a core capability of language models (LMs), yet it remains unclear how much model capacity is necessary to support reasoning during pretraining.
arXiv:2502. 15543v4 Announce Type: replace-cross Abstract: Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external evidence.
arXiv:2601.07794v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly evaluated on their ability to perform multi-hop reasoning, i.e., to combine multiple pieces of...
arXiv:2608. 04519v1 Announce Type: new Abstract: Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not.
arXiv:2602. 02470v2 Announce Type: replace Abstract: Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" -- when trained on forward knowledge data of the form "$A \rightarrow B$" (e.
arXiv:2509. 24653v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel at multi-hop reasoning in distribution, yet fail on unseen compositions, a phenomenon known as the curse of two-hop reasoning.