Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this d...
The paper investigates how sequential knowledge editing can degrade a language model’s ability to discern reliable evidence from unreliable evidence without affecting overall accuracy. Using a conservatively tuned LoRA on Qwen2.5‑7B‑Instruct, the authors show that after 1,000 edits the model’s arbitration score for untouched facts drops by 36%, leading to higher error rates on its most confident decisions, while MMLU accuracy remains unchanged. The study also finds that in some model‑method combinations, sequential edits can reduce MMLU to chance levels even though edit success and locality remain perfect.
By Atul Anand
The paper reports that existing knowledge‑editing benchmarks cannot evaluate the scope decision—whether a stored edit applies to a query—because they are counterfactual and lack negative examples. Using the gradient‑free editor INLAY, the authors exhaustively test every router action on 1,689 queries across three datasets and find that an oracle router achieves no gain over a static policy, and abstention never wins. The authors attribute this to the structural design of the benchmarks and demonstrate that adding a missing negative condition restores some headroom and allows abstention to win.
By Aditya Pratap Singh
The paper argues that evaluating continual knowledge‑updating methods solely at a final checkpoint and a single adapter rank can be misleading. By fixing a periodic hierarchy and comparing it to cumulative replay on a 24‑month Wikidata stream, the authors show that the apparent best method changes depending on the evaluation month, replay LoRA rank, and query formulation. They recommend reporting performance trajectories and capacity sweeps, and only declaring a robust winner when the ranking remains stable across the evaluation region.
By Heejin Choi
arXiv:2609. 22880v1 Announce Type: cross Abstract: LLM rerankers add of the order of \$0.
By Andre Bacellar
arXiv:2608. 09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one.
By Rose Cymbler, Daniel Guez, Laurent Fabre
arXiv:2608.21610v1 Announce Type: new
Abstract: Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform chunks and retrieved by similarity, dis...
By Junaid Farooq
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious...
arXiv:2608. 20106v1 Announce Type: new Abstract: We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers.
By Nikita Khudov
The paper investigates time‑incremental continued pretraining (CPT) of large language models using web‑scale data that overlaps across snapshots. Across six open‑weight models, CPT improves factual recall without catastrophic forgetting, while the cost is negligible and data quality outweighs quantity. The study identifies optimal learning rates, shows LoRA can match full CPT, and demonstrates that CPT gains transfer to fine‑tuned models.
By F{\i}rat \"Oncel, Salman Hussain Ali, Mirco Ravanelli, Cem Subakan, \c{C}a\u{g}atay Y{\i}ld{\i}z
Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.
By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
arXiv:2607. 12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets.
By Taizhen Cheung, SA Kwon