arXiv:2606. 18509v1 Announce Type: new Abstract: Reliable generalization in conditional latent variable models requires understanding both identifiability and extrapolation: how observed variation across attributes determines latent structure, and how that structure determines distributions at unseen attributes.
By Soheun Yi, Yizhou Lu, Chandler Squires, Pradeep Ravikumar
arXiv:2606. 04045v1 Announce Type: cross Abstract: Representation learning is often described as preserving the information in an input that is relevant for prediction.
By Vasileios Sevetlidis
The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.
By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran
arXiv:2511. 22823v2 Announce Type: replace-cross Abstract: Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly or infeasible to acquire.
By Miao Zhang, Junpeng Li, Changchun Hua, Yana Yang
arXiv:2503. 21796v2 Announce Type: replace-cross Abstract: Self-supervised learning has become an increasingly important paradigm in the domain of machine intelligence.
By Alexander Ororbia, Karl Friston, Rajesh P. N. Rao
arXiv:2607. 00006v1 Announce Type: cross Abstract: Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering.
By Shuaizhi Cheng
arXiv:2609.06546v1 Announce Type: cross
Abstract: Although machine learning can be used to predict the evolution of physical systems from data, a formulation that learns only the system state at each...
By Sai Siddharth, Maniarasu Ravi
The paper proposes treating the ‘unit’—a persistent referent that multiple events may refer to—as an explicit primitive in machine learning tasks. It formalizes supervised learning as learning a pair of a tokenizer that generates a contextual unit token and a shared response law that uses this token, thereby distinguishing homogeneous from heterogeneous worlds. The work also introduces concepts such as unit abduction and trusted resolvers to handle cases where unit identity is unresolved.
By Heyang Gong
arXiv:2606. 08390v1 Announce Type: new Abstract: When a neural time-series model reports that one variable modulates another's effect on a target, is the discovered interaction a property of the data or an artifact of model flexibility?
By Valentina Kuskova, Dmitry Zaytsev, Michael Coppedge
Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.
By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts
arXiv:2608. 01383v1 Announce Type: new Abstract: Masked prediction learns representations by fitting a schedule-weighted family of conditional laws, but it remains unclear when near-optimal conditional prediction pins down the underlying joint law.
By Yichao Cai, Javen Qinfeng Shi
arXiv:2502.04131v2 Announce Type: replace
Abstract: The successful application of modern machine learning for time series classification is often hampered by limitations in quality and quantity of av...
By Janis Norden, Elisa Oostwal, Michael Chappell, Peter Tino, Kerstin Bunte