arXiv AI

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

arXiv:2608. 12218v1 Announce Type: cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories.

arXiv AI
Sep 4

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

The paper investigates how large language models acquire knowledge during pre‑training, proposing that auxiliary views—reformulations of knowledge—are causally beneficial. Experiments show that repetition is essential, paraphrasing helps only at smaller batch sizes, and reallocating tokens from repetition to auxiliary views improves learning even for factual recall. The study also finds that the benefit of auxiliary views does not depend on the teacher model’s strength, identifies specific knowledge types that aid learning, and explores mechanistic effects via layer‑wise biases and compression.

By Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen
Hugging Face Trending Papers
Jul 2

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization.

arXiv AI
Jul 8

LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

arXiv:2607. 06160v1 Announce Type: cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision.

By Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu
arXiv AI
6d ago

MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries

MoSAR introduces a mixture of semantic attention regimes that learns an adaptive, distance‑dependent attention geometry from data, rather than predefining sparse or local patterns. The model uses input‑conditioned routers to select short, medium, or global regimes, creating a continuous attention field that can be discretized for efficient inference. Experiments show that MoSAR achieves lower‑reach attention without sacrificing language‑modeling quality, improving perplexity over dense RoPE and outperforming baselines like ALiBi, while remaining stable under top‑1 discretization.

By Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos