arXiv:2608. 15976v1 Announce Type: new Abstract: A learning system can occupy execution states that are indistinguishable under every declared present-behavior readout yet respond differently to future training.
By Qinyou Wang
arXiv:2608. 10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it.
By Ashim Dhor, Pin-Yu Chen
Revelation Control studies how to price interventions that reveal hidden state only when the revealed distinctions can alter a consequential decision, while separately accounting for any useful progress the intervention itself creates. The authors develop a framework for learning systems that defines decision‑sufficient revelation, revelation depth, and a cost‑adjusted factorization criterion, and they provide a target‑independent protocol for model‑specific instantiation. Experiments on Qwen2.5‑7B and Mistral‑7B‑v0.3 show that deeper future‑learning probes have positive decision value and that productive reuse yields strict equal‑compute utility advantages, supporting a structural transfer of the decision theory and evaluation protocol across architectures.
By Qinyou Wang
arXiv:2607. 04333v1 Announce Type: new Abstract: Grokking -- generalization arriving long after training-set interpolation -- can be accelerated by structure-agnostic interventions: gradient filtering, weight-norm clamping, geometric penalties on hidden representations.
By Gunner Levi Howe
arXiv:2609.36375v1 Announce Type: new
Abstract: Continual learning is usually studied through mechanisms that preserve old knowledge. We develop Successional Learning Theory (SLT), a mesoscopic accou...
By Shoaib Ahmed Dipu, Md Salman Shamil, Sayeed Shafayet Chowdhury
arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.
By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
arXiv:2609. 23366v1 Announce Type: new Abstract: Recurrent models must preserve information that changes future behavior while suppressing hidden-state error.
By Linzhe Zhang, Changming Xu
arXiv:2512. 18471v2 Announce Type: replace Abstract: Continual learning systems face a fundamental geometric obstacle: as experience accumulates on a fixed-capacity manifold, covering numbers grow linearly with time, eventually forcing representational overlap and catastrophic interference.
By Xin Li
ForeTime‑VLA is a causal vision‑language‑action policy that distills future‑aware representations from a frozen Fast‑WAM teacher, enabling it to anticipate contact events during conveyor‑belt manipulation. The method compresses current and future video latents into a 64‑dimensional target, uses an eight‑frame history encoder to predict this target along with manipulation phase and time‑to‑transition, and conditions a VLM prefix on future tokens and phase. On a deduplicated conveyor‑belt dataset, ForeTime‑VLA reduces test MAE by 2.63% and L2 by 3.02%, while real‑robot experiments show significantly higher grasp success rates compared to the next‑best reference.
whyItMatters":"The approach demonstrates that distilling future‑token knowledge from a world‑action model can improve dynamic manipulation performance without the computational cost of running the teacher at inference time."
By Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang
The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, defining it via a continuation‑conditioned value of a constrained learning‑state intervention and observation‑relative fibers. It presents an executable finite‑frame classifier‑head that protects current logits while repairing historical margins, and distinguishes local admissibility, intervention value, and complete‑policy performance. Experiments on CLINC‑derived class‑incremental tasks, output distillation with RoBERTa, and SGDW dynamics demonstrate that coupling can produce both positive and negative interactions, and that coordinated content controls can match or exceed development gains while guided allocation reduces cross‑entropy loss compared to standard replay.
By Qinyou Wang
arXiv:2607. 25387v2 Announce Type: replace Abstract: Learned restriction maps in sheaf graph neural networks are often treated as proof that the model has discovered useful edge geometry.
By Yi Liu
arXiv:2606. 05957v1 Announce Type: new Abstract: Singular learning theory and information geometry have studied the same parameter spaces in mostly separate vocabularies: the former computes Bayesian invariants in resolved coordinates, the latter works in original coordinates under a non-degeneracy assumption that overparameterised models routinely violate.
By Tejas Pradeep Shirodkar