arXiv Machine Learning By Beatrix M. G. Nielsen, Emanuele Marconato, Luigi Gresele, Andrea Dittadi, Simon Buchholz

Logit Distance Bounds Representational Similarity

Read the original on arXiv Machine Learning →

arXiv:2602. 15438v3 Announce Type: replace Abstract: For a broad family of discriminative models that includes autoregressive language models, identifiability results imply that if two models induce the same conditional distributions, then their internal representations are equal up to an invertible linear transformation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.

By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
arXiv AI
6d ago

A Flow Matching Framework for Neural Representational Dissimilarity

The paper introduces a flow matching framework that unifies various neural representational dissimilarity metrics under a single theoretical umbrella. By interpreting these metrics as Jeffreys divergences with different velocity constraints, the authors demonstrate that flow matching improves distance estimation for complex distributions and continuous variables. The framework also facilitates the principled design of new dissimilarity measures.

By Zeyuan Ye, Xue-Xin Wei