MMLA: How Memory Lets the Past Shape the Future
arXiv:2606. 28876v3 Announce Type: replace-cross Abstract: Proposal.
arXiv:2606. 28876v3 Announce Type: replace-cross Abstract: Proposal.
arXiv:2608.20873v1 Announce Type: new Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint,...
arXiv:2608. 03137v1 Announce Type: new Abstract: Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction.
arXiv:2607. 12204v2 Announce Type: replace Abstract: Auditable memory requires a precise contract: which output is preserved, relative to which reference solve, and across which updates.
arXiv:2606. 10616v1 Announce Type: new Abstract: Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their finite context windows, making memory retention a fundamental resource-allocation problem.
arXiv:2607. 27539v2 Announce Type: replace Abstract: Exact deletion from persistent language-model memory depends on whether a record's effect remains addressable after later computation.
arXiv:2608.20965v1 Announce Type: new Abstract: We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and rel...
arXiv:2609.38142v1 Announce Type: new Abstract: A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advi...
UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."
arXiv:2609.14138v1 Announce Type: cross Abstract: As LLM agents become integrated into increasingly complex workflows, they must continually acquire new capabilities while retaining competence on pre...
The paper introduces Boundary‑Calibrated Intervention Transfer (BCIT), a method for conditional experience transfer in autonomous large language model (LLM) post‑training. BCIT links each past update to its specific parent model, data, and training stage, checks whether those conditions still hold, vetoes updates with hard conflicts, and, when necessary, runs a bounded training trial to confirm applicability before adopting the update. Experiments on a 4B model across finance reasoning, text‑to‑SQL, and function calling show that BCIT reduces harmful updates and achieves higher final‑model quality under equal computational budgets compared to other approaches.
arXiv:2609.25853v1 Announce Type: new Abstract: Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be...