arXiv AI By Guangsheng Yu, Yanna Jiang, Qin Wang, Baihe Ma, Xu Wang

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

Read the original on arXiv AI →

K-Bench is a new benchmark designed to evaluate large language model (LLM) unlearning when the models are deployed as agents. Unlike previous benchmarks that only inspect the final answer, K-Bench examines all six channels of a ReAct agent—including chain-of-thought, tool calls, tool observations, and elicited summaries—to determine if a secret is leaked. The benchmark measures leakage for secrets placed in the model weights, prompt, or retrieval store, and finds that many existing unlearning methods fail to prevent leaks in deployed agents, especially when secrets reside in the prompt or retrieval store.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 16

Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations

arXiv:2606. 15903v1 Announce Type: cross Abstract: Where an LLM sits in an agent memory pipeline -- between the recall plane that retrieves stored facts (extensively benchmarked) and the control plane that mutates them via supersede, release, purge (largely untested) -- shapes which forgetting failure modes the system recovers.

By Dongxu Yang
arXiv AI
Aug 25

Repo2Skill-Evo: Repository Skills Go Stale in Silence

arXiv:2608.21964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural kno...

By Chenyuan Duan, Ge Shi, Zineng Mao, Ge Zhang, Hao Liang, Yinzhu Piao, Yuchen Wu, Zhixin Yao, Kaiyu Huang, Wenhao Huang, Linzhuang Sun, Shen Yan, Wentao Zhang
arXiv AI
Jun 10

Deployment-Time Memorization in Foundation-Model Agents

arXiv:2606. 10062v1 Announce Type: new Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time function rather than solely a property of model weights.

By Lei (Rachel), Chen, Guilin Zhang, Kai Zhao, Dalmo Cirne, Andy Olsen, Xu Chu, Zeke Miller, Alet Blanken, Amine Anoun, Jerry Ting