arXiv AI

A Fiber Criterion for Representation Identifiability in Supervised Learning

arXiv:2606. 01092v1 Announce Type: cross Abstract: Supervised learning evaluates predictors through their input-output behavior.

arXiv Machine Learning
Jun 18

Concept Modulation Models: A Unified Framework for Identifiability and Extrapolation

arXiv:2606. 18509v1 Announce Type: new Abstract: Reliable generalization in conditional latent variable models requires understanding both identifiability and extrapolation: how observed variation across attributes determines latent structure, and how that structure determines distributions at unseen attributes.

By Soheun Yi, Yizhou Lu, Chandler Squires, Pradeep Ravikumar
arXiv Machine Learning
1d ago

Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions

The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.

By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran
arXiv AI
Jul 2

Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem

arXiv:2607. 00006v1 Announce Type: cross Abstract: Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering.

By Shuaizhi Cheng
arXiv Machine Learning
Sep 10

Introductory Notes on Learning$^2$

arXiv:2609.06546v1 Announce Type: cross Abstract: Although machine learning can be used to predict the evolution of physical systems from data, a formulation that learns only the system state at each...

By Sai Siddharth, Maniarasu Ravi
arXiv Machine Learning
Aug 27

Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events

The paper proposes treating the ‘unit’—a persistent referent that multiple events may refer to—as an explicit primitive in machine learning tasks. It formalizes supervised learning as learning a pair of a tokenizer that generates a contextual unit token and a shared response law that uses this token, thereby distinguishing homogeneous from heterogeneous worlds. The work also introduces concepts such as unit abduction and trusted resolvers to handle cases where unit identity is unresolved.

By Heyang Gong
arXiv Machine Learning
Sep 23

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.

By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts