The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, defining it via a continuation‑conditioned value of a constrained learning‑state intervention and observation‑relative fibers. It presents an executable finite‑frame classifier‑head that protects current logits while repairing historical margins, and distinguishes local admissibility, intervention value, and complete‑policy performance. Experiments on CLINC‑derived class‑incremental tasks, output distillation with RoBERTa, and SGDW dynamics demonstrate that coupling can produce both positive and negative interactions, and that coordinated content controls can match or exceed development gains while guided allocation reduces cross‑entropy loss compared to standard replay.
By Qinyou Wang
arXiv:2608.20965v1 Announce Type: new
Abstract: We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and rel...
By Mian Wang
arXiv:2608. 11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting.
By Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu
arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.
By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv:2607. 14185v1 Announce Type: cross Abstract: Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback.
By Xuening Wu, Shan Yu, Shenqin Yin
arXiv:2602. 10430v2 Announce Type: replace-cross Abstract: Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures.
By Yusen Huo, Changping Wang, Yangru Huang, Jun Zhang, Jie Jiang