arXiv AI

Benchmarking Knowledge Editing using Logical Rules

arXiv:2606. 10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge.

arXiv AI
Aug 13

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

arXiv:2608. 11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world.

By Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao
arXiv Computation and Language
Aug 27

Addressing the Reasoning Gap: Mechanistic Circuit-Based Knowledge Editing in Large Language Models

The paper introduces MCircKE, a mechanistic circuit-based knowledge editing framework for large language models. MCircKE identifies the causal circuits involved in a specific reasoning task and surgically updates parameters only within those circuits, thereby addressing the reasoning gap where edited facts are not used in multi-step reasoning. Experiments on the MQuAKE-series benchmarks show that this approach improves multi-hop reasoning performance after knowledge editing.

By Tianyi Zhao, Yinhan He, Wendy Zheng, Chen Chen
Hugging Face Trending Papers
Jul 29

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification.

arXiv Computation and Language
Sep 2

InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing

InComeS is a framework that improves large language models (LLMs) for model editing by compressing editing contexts into a key‑value cache associated with a special gist token. It adds cross‑attention modules to dynamically select the most relevant information from these gist pools, allowing efficient handling of multiple edits beyond the LLM’s context window. Experiments on diverse editing benchmarks show that InComeS enhances both effectiveness and efficiency compared to existing methods.

By Shuaiyi Li, Zhisong Zhang, Yang Deng, Chenlong Deng, Tianqing Fang, Hongming Zhang, Haitao Mi, Dong Yu, Wai Lam
arXiv Computation and Language
Aug 25

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

arXiv:2608.22753v1 Announce Type: new Abstract: Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided proced...

By Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu