arXiv Machine Learning

MMLA: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation

arXiv AI
Jun 10

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

arXiv:2606. 10616v1 Announce Type: new Abstract: Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their finite context windows, making memory retention a fundamental resource-allocation problem.

By Qingcan Kang, Liu Mingyang, Shixiong Kai, Kaichao Liang, Tao Zhong, Mingxuan Yuan
arXiv AI
Aug 24

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations. "whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."

By Ye Chen, Weining Zhang
arXiv AI
Aug 28

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

The paper introduces Boundary‑Calibrated Intervention Transfer (BCIT), a method for conditional experience transfer in autonomous large language model (LLM) post‑training. BCIT links each past update to its specific parent model, data, and training stage, checks whether those conditions still hold, vetoes updates with hard conflicts, and, when necessary, runs a bounded training trial to confirm applicability before adopting the update. Experiments on a 4B model across finance reasoning, text‑to‑SQL, and function calling show that BCIT reduces harmful updates and achieves higher final‑model quality under equal computational budgets compared to other approaches.

By Tingyun Li, Wenfeng Feng, Weiqing Li, Abudukelimu Wuerkaixi, Guohua Liu, Yuewei Zhang