arXiv AI By Ya Gao, Kalle Kujanp\"a\"a, Pekka Marttinen, Harri Valpola, Alexander Ilin

Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories

Read the original on arXiv AI →

arXiv:2602. 02028v2 Announce Type: replace Abstract: Enabling artificial intelligence systems, particularly large language models, to update knowledge and flexibly apply it during reasoning remains a central challenge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

Addressing the Reasoning Gap: Mechanistic Circuit-Based Knowledge Editing in Large Language Models

The paper introduces MCircKE, a mechanistic circuit-based knowledge editing framework for large language models. MCircKE identifies the causal circuits involved in a specific reasoning task and surgically updates parameters only within those circuits, thereby addressing the reasoning gap where edited facts are not used in multi-step reasoning. Experiments on the MQuAKE-series benchmarks show that this approach improves multi-hop reasoning performance after knowledge editing.

By Tianyi Zhao, Yinhan He, Wendy Zheng, Chen Chen
arXiv Computation and Language
1d ago

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

EngramEdit introduces a method for decoupled knowledge updates in large language models using conditional memory architectures. By computing target memory representations that align with updated facts across multiple expressions, it jointly adjusts shared n‑gram embeddings while penalizing frequent ones to preserve unrelated knowledge. Experiments demonstrate near‑perfect editing success, improved multi‑hop reasoning, and strong performance under chain‑of‑thought prompting, all while maintaining general capabilities.

By Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu, Ning Song, Yongqi Li, Wenjie Li
arXiv Machine Learning
Sep 23

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Ladders-of-Thought (LoT) is a framework that enhances reasoning in small- to mid-scale large language models by automatically generating easier variants of reasoning problems and organizing them into difficulty buckets. It uses a self‑evolving bandit scheduler to adaptively allocate training, improving performance across math and multi‑hop reasoning tasks on 1–8 B models. LoT achieves significant gains (e.g., +32 pp on AddSub, +16 pp on QASC) and converges faster than staged curricula.

By Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang
arXiv Machine Learning
Aug 27

From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

The paper introduces Golden-GRPO Injection (GRIN), a three-stage self‑learning framework that uses a mixed‑policy reinforcement learning algorithm to inject knowledge into large language models. GRIN injects a golden answer to provide learning signals even when on‑policy rollouts fail on novel facts, and is evaluated on two new document‑level benchmarks—Blank and Counter—that test novel acquisition and counterfactual overwrite. Experiments show that mixed‑policy RL enables knowledge absorption beyond what supervised fine‑tuning can achieve, with GRIN outperforming SFT and other RL baselines on harder question types while matching them on basic fact recall.

By Zhibo Hou, Fan Zhao, Zhiyu An, Wan Du
arXiv AI
Aug 13

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

arXiv:2608. 11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world.

By Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao