The paper investigates whether computation from earlier problems can aid large language models (LLMs) in solving subsequent ones. Preliminary experiments show that retained conversation history can both improve and degrade later-turn accuracy. To address this, the authors propose STAIR, a lightweight module that stores key-value pairs from previous responses and redirects new queries to this bank, improving later-turn accuracy by up to 11.67 percentage points across several benchmarks.
By Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song
arXiv:2602. 24287v2 Announce Type: replace-cross Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses.
By Jenny Y. Huang, Leshem Choshen, Wei Sun, Omar Khattab, Ram\'on Fernandez Astudillo, Mehul Damani, Tamara Broderick, Jacob Andreas
RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.
By Yuan Si, Simeng Han, Daming Li, Jialu Zhang
arXiv:2607. 01935v1 Announce Type: new Abstract: Long term memory lets LLM agents act as persistent assistants, but user facts change.
By Zitong Shi, Yixuan Tang, Anthony Kum Hoe Tung
Grounded Continuation introduces a runtime verifier that classifies each utterance in an LLM conversation into one of eight epistemic operations and uses a symbolic engine to maintain a dependency map of claims and their supports. The verifier checks whether a new continuation is grounded by walking this map, a linear-time process that requires no additional LLM calls. On benchmarks such as ReviseQA and MemoryAgentBench, the verifier improves single-hop accuracy for several QA models, even enabling a 7B model to outperform GPT‑4o when guided by the verifier.
By Qisong He, Jinwei Hu, Xinmiao Huang, Changshun Wu, Yi Dong, Xiaowei Huang
arXiv:2608. 08300v1 Announce Type: new Abstract: Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions.
By Hakeem Hannoon, Andrew Zhao, Mihir Narayan, Sharvin Goyal, Ivaxi Sheth
arXiv:2609.07093v2 Announce Type: replace
Abstract: Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely a...
By Yifan Wang, Xinkui Lin, Yongxiu Xu, Shen Gao, Ruochen Yang, Kun Huang, Yubin Wang, Jie Wu, Wei Liu, Jian Luan, Hongbo Xu, Shuo Shang
The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.
By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv:2608. 06296v1 Announce Type: new Abstract: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs).
By Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos
arXiv:2609.05882v1 Announce Type: cross
Abstract: Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantia...
By Jinnan Li, Zheren Fu, Yue Wang, Jinzhe Li, Yuan Wu, Yi Chang
arXiv:2601. 00821v4 Announce Type: replace Abstract: A growing class of conversational-memory systems compresses dialogue history into structured artifacts (extracted facts, decisions, or events) on the premise that distilled structure retrieves better than raw text.
By Tao An
The paper identifies a specific issue in supervised fine‑tuning (SFT) of large language models called factual access failure, where models can recognize correct facts under constrained tests but fail to generate them in open‑ended settings. It demonstrates that SFT can cause both genuine wrong answers and expression‑level errors such as verbosity or formatting mismatches. To mitigate this, the authors propose Recall‑Anchored Distillation (RAD), a self‑distillation method that aligns the fine‑tuned model with the base model’s soft output distribution on unlabeled out‑of‑distribution text, thereby recovering lost factual recall without needing labeled data.
By Haodong Chen, Yadong Wang, Shengtao Wen, Dong Liang, Xiang Chen