arXiv:2608.01383v3 Announce Type: replace
Abstract: Masked prediction learns by inferring missing variables from visible context. This raises a fundamental question: when does near-optimal conditiona...
By Yichao Cai, Javen Qinfeng Shi
The paper investigates the limits of the maximal coding rate reduction (MCR²) framework for out‑of‑distribution (OOD) generalisation. It shows that MCR² can lead to complete prediction failure under distribution shift, even when a perfectly stable feature is available, and that adding invariance principles from IRM or REx does not resolve this issue. The authors conclude that additional assumptions or learning principles are needed to guarantee stable OOD predictions with MCR².
By Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi, Po Yang
arXiv:2607. 18305v1 Announce Type: cross Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text.
By Priyansh Srivastava, Romit Chatterjee
arXiv:2610. 01578v1 Announce Type: new Abstract: Why can masked prediction learn useful representations that unmasked reconstruction misses?
By Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborov\'a
The paper introduces the concept of task‑conditioned active observability, defining the minimal interaction cost needed for an autonomous agent to identify task‑relevant states while guaranteeing safe abstention. It formalizes this complexity, proving that task‑predictive equivalence yields a unique minimal sufficient quotient that preserves complexity and eliminates unnecessary distinctions. The authors present theoretical characterizations for deterministic and noisy regimes, and demonstrate a certified observer that reduces sensor usage and model steps while maintaining zero false acceptances in extensive high‑dimensional trials.
By Linzhe Zhang, Changming Xu
arXiv:2608. 10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it.
By Ashim Dhor, Pin-Yu Chen