arXiv Machine Learning

Streaming Knowledge Compilation: Proactive Materiality-Scored Pinning for Time-Evolving LLM Wikis

arXiv:2606. 09877v1 Announce Type: new Abstract: LLM wiki systems compile knowledge into pre-filled KV caches for efficient inference, but assume a static corpus -- an assumption that fails whenever the underlying information landscape evolves.

arXiv AI
Sep 25

Sequential knowledge editing breaks a model's ability to tell good evidence from bad, without costing it accuracy

The paper investigates how sequential knowledge editing can degrade a language model’s ability to discern reliable evidence from unreliable evidence without affecting overall accuracy. Using a conservatively tuned LoRA on Qwen2.5‑7B‑Instruct, the authors show that after 1,000 edits the model’s arbitration score for untouched facts drops by 36%, leading to higher error rates on its most confident decisions, while MMLU accuracy remains unchanged. The study also finds that in some model‑method combinations, sequential edits can reduce MMLU to chance levels even though edit success and locality remain perfect.

By Atul Anand
arXiv AI
Aug 28

On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study

The paper reports that existing knowledge‑editing benchmarks cannot evaluate the scope decision—whether a stored edit applies to a query—because they are counterfactual and lack negative examples. Using the gradient‑free editor INLAY, the authors exhaustively test every router action on 1,689 queries across three datasets and find that an oracle router achieves no gain over a static policy, and abstention never wins. The authors attribute this to the structural design of the benchmarks and demonstrate that adding a missing negative condition restores some headroom and allows abstention to win.

By Aditya Pratap Singh
arXiv Machine Learning
Sep 4

Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating

The paper argues that evaluating continual knowledge‑updating methods solely at a final checkpoint and a single adapter rank can be misleading. By fixing a periodic hierarchy and comparing it to cumulative replay on a 24‑month Wikidata stream, the authors show that the apparent best method changes depending on the evaluation month, replay LoRA rank, and query formulation. They recommend reporting performance trajectories and capacity sweeps, and only declaring a robust winner when the ranking remains stable across the evaluation region.

By Heejin Choi
arXiv Computation and Language
Sep 22

Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

The paper investigates time‑incremental continued pretraining (CPT) of large language models using web‑scale data that overlaps across snapshots. Across six open‑weight models, CPT improves factual recall without catastrophic forgetting, while the cost is negligible and data quality outweighs quantity. The study identifies optimal learning rates, shows LoRA can match full CPT, and demonstrates that CPT gains transfer to fine‑tuned models.

By F{\i}rat \"Oncel, Salman Hussain Ali, Mirco Ravanelli, Cem Subakan, \c{C}a\u{g}atay Y{\i}ld{\i}z
arXiv AI
Aug 24

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.

By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen