arXiv AI By Xuening Wu, Shan Yu, Shenqin Yin

Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape

Read the original on arXiv AI →

arXiv:2607. 14185v1 Announce Type: cross Abstract: Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 24

Intrinsic-Extrinsic Coupling in Learning Dynamics

The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, where a learner’s current observations do not solely dictate its future training responses. It formalizes this coupling through a continuation‑conditioned value of a constrained learning‑state intervention and employs an executable finite‑frame classifier‑head to protect current logits while adjusting historical margins. Experiments across CLINC‑derived class‑incremental settings, output distillation with RoBERTa, and SGDW dynamics reveal that coupling can produce both positive and negative interactions, and that coordinated interventions can match or exceed development gains while reducing cross‑entropy loss compared to standard replay.

arXiv Machine Learning
Sep 25

Intrinsic-Extrinsic Coupling in Learning Dynamics

The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, defining it via a continuation‑conditioned value of a constrained learning‑state intervention and observation‑relative fibers. It presents an executable finite‑frame classifier‑head that protects current logits while repairing historical margins, and distinguishes local admissibility, intervention value, and complete‑policy performance. Experiments on CLINC‑derived class‑incremental tasks, output distillation with RoBERTa, and SGDW dynamics demonstrate that coupling can produce both positive and negative interactions, and that coordinated content controls can match or exceed development gains while guided allocation reduces cross‑entropy loss compared to standard replay.

By Qinyou Wang
arXiv AI
Jul 7

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.

By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi
arXiv Machine Learning
Aug 28

Stable but Wrong: When Learning Stabilizes Away from the Truth

The paper introduces the concept of Stable but Wrong (SBW), describing situations where a learning process appears stable yet converges to a solution that is systematically biased away from a true objective. Using a minimal strongly convex model, the authors demonstrate that persistent bias in update directions can shift the convergence point from the optimal solution. Experiments across reinforcement learning, supervised learning, and continual fine-tuning of large language models reveal a recurring disconnect between normal optimization behavior and correctness under both static and feedback‑coupled biases, and show that recovery interventions can still modify subsequent learning trajectories.

By Zhipeng Zhang