arXiv AI By Arda Uzunoglu, Benjamin van Durme, Daniel Khashabi

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

Read the original on arXiv AI →

arXiv:2608. 12218v1 Announce Type: cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

The paper investigates how large language models acquire knowledge during pre‑training, proposing that auxiliary views—reformulations of knowledge—are causally beneficial. Experiments show that repetition is essential, paraphrasing helps only at smaller batch sizes, and reallocating tokens from repetition to auxiliary views improves learning even for factual recall. The study also finds that the benefit of auxiliary views does not depend on the teacher model’s strength, identifies specific knowledge types that aid learning, and explores mechanistic effects via layer‑wise biases and compression.

By Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen