arXiv Machine Learning

Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling

arXiv Machine Learning
4d ago

SCAMP: Sparse-anchor Control is One Small Projection

SCAMP introduces a training‑free, damped Gauss‑Newton method that adjusts only the sparse anchor points in a frozen differentiable decoder, keeping the rest of the state unchanged. By operating solely in the space of the anchors’ Jacobian rows, it solves a system whose size matches the number of anchor constraints rather than the full state, enabling efficient control across diverse text‑to‑motion generators. Applied to seven existing generators, SCAMP achieves anchor errors that match or surpass all released control methods and can close anchors on hosts that originally lacked them.

By Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan, Yanwen Guo, Hansung Kim, Xiaohao Cai, Dongjie Fu
arXiv Machine Learning
Jul 1

Predictable GRPO: A Closed-Form Model of Training Dynamics

arXiv:2606. 30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error.

By Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta
arXiv Machine Learning
1d ago

Stable and Counterfactually Robust Physical World Models from Imposed Structure and Learned Physics

The paper introduces a world model that learns to predict the evolution of physical systems while respecting key physical principles. By hard‑coding a general structure—generating dynamics from the gradient of a learned energy via a fixed reversible operator and imposing constraints on energy, dissipation, and interventions—the model achieves second‑law compatible dissipation, accurate responses to parameter changes, long‑term stability, and robustness to disturbances. Experiments on an electromagnetic cavity, a particle‑in‑cell grid, and shallow‑water fluid demonstrate that the model can recover accurate constitutive functions, distinguish conserving from dissipating regimes, and transfer learned physics to unseen conditions, outperforming unconstrained models.

By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv AI
Jun 12

Order Is Not Control

arXiv:2606. 12923v1 Announce Type: cross Abstract: AI alignment, interpretability, steering, and neural perturbation studies identify order-inducing objects.

By Gareth Seneque, Lap-Hang Ho, Nafise Erfanian Saeedi, Jeffrey Molendijk, Tim Elson
arXiv Machine Learning
Sep 25

Intrinsic-Extrinsic Coupling in Learning Dynamics

The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, defining it via a continuation‑conditioned value of a constrained learning‑state intervention and observation‑relative fibers. It presents an executable finite‑frame classifier‑head that protects current logits while repairing historical margins, and distinguishes local admissibility, intervention value, and complete‑policy performance. Experiments on CLINC‑derived class‑incremental tasks, output distillation with RoBERTa, and SGDW dynamics demonstrate that coupling can produce both positive and negative interactions, and that coordinated content controls can match or exceed development gains while guided allocation reduces cross‑entropy loss compared to standard replay.

By Qinyou Wang
Hugging Face Trending Papers
Sep 24

Intrinsic-Extrinsic Coupling in Learning Dynamics

The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, where a learner’s current observations do not solely dictate its future training responses. It formalizes this coupling through a continuation‑conditioned value of a constrained learning‑state intervention and employs an executable finite‑frame classifier‑head to protect current logits while adjusting historical margins. Experiments across CLINC‑derived class‑incremental settings, output distillation with RoBERTa, and SGDW dynamics reveal that coupling can produce both positive and negative interactions, and that coordinated interventions can match or exceed development gains while reducing cross‑entropy loss compared to standard replay.

arXiv AI
2d ago

The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State

The paper investigates how inherited state affects sub-agent performance in multi-agent frameworks, comparing three inheritance policies—Reset, Selective, and Full—across a ladder of Qwen3 models. It finds that reliance on stale state decreases with model capability, but a mid-capability model (Qwen3‑1.7B) exhibits a statistically significant local minimum of net harm, defining a "danger band." Selective handoff consistently improves accuracy over Full, especially within the danger band, while a fixed-threshold router fails on other datasets.

By Jundong Hu, Shekar Ramachandran
arXiv AI
Jul 15

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

arXiv:2607. 12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent.

By Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang
arXiv Machine Learning
Sep 17

The Automaton Underneath: The Additive Input Pathway Is a Parasitic Attractor for State Tracking in Householder Linear RNN

arXiv:2609. 18966v1 Announce Type: new Abstract: Linear RNNs with input-dependent Householder-product transitions (DeltaNet/DeltaProduct-class) can provably represent hard state-tracking automata, yet trained models fail to length-generalize -- a gap recent work attributes to optimization, without a causal account.

By Gunner Levi Howe