Multi-Hop Knowledge Composition is Bound by Pretraining Exposure
Large Language Models fail at implicit multi-hop reasoning: a model answers "When was $X$ born? " and "Who is $Y$'s closest friend?
Large Language Models fail at implicit multi-hop reasoning: a model answers "When was $X$ born? " and "Who is $Y$'s closest friend?
arXiv:2609.34187v2 Announce Type: replace-cross Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
Large Language Models struggle with implicit multi‑hop reasoning, correctly answering individual facts but failing to combine them in a single pass. In a controlled setting, the authors show that this failure persists even with high 1‑hop accuracy, indicating it is due to pretraining exposure rather than missing knowledge. They test nine data‑centric augmentation formats and find that only individuals seen in compositional contexts during pretraining enable transfer to unseen questions, proving exposure to such contexts is necessary for implicit multi‑hop reasoning.
arXiv:2606. 17945v1 Announce Type: new Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered.
arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...
arXiv:2607. 16097v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it.
Large language models can solve complex multi‑hop tasks but often fail on simple two‑hop queries, even when each hop is individually correct. In a controlled symbolic setting, the authors find that models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Mechanistic analysis shows that successful generalization relies on consistent intermediate representations across contexts, whereas failures arise from a mismatch between lower‑layer representation construction and upper‑layer mapping to outputs. The study proposes a recurrent‑style training strategy that improves out‑of‑distribution two‑hop generalization.
arXiv:2607. 19345v1 Announce Type: cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier.
arXiv:2608.30627v1 Announce Type: new Abstract: As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token pred...
arXiv:2410.17021v2 Announce Type: replace Abstract: Large Language Models with chain-of-thought prompting, such as OpenAI-o1, have shown impressive capabilities in natural language inference tasks. H...
arXiv:2608. 03930v1 Announce Type: cross Abstract: Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition.
arXiv:2610.00673v1 Announce Type: cross Abstract: Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage...