arXiv Machine Learning

More Data Cannot Break a Symmetry: Identifiability by Design

The paper shows that unsupervised representational alignment can fail due to symmetry in the stimulus geometry, even before data are collected. By using a design-time diagnostic based on the automorphism group of the geometry, the authors demonstrate that dense sampling can create near-duplicates that make alignment degenerate. Applying this diagnostic to a colour design reduces catastrophic alignment failures from 75% to 2% without altering models, layers, or solvers.

arXiv Machine Learning
Jul 8

Geometric Stability: The Missing Axis of Representations

arXiv:2601. 09173v5 Announce Type: replace Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is reliably recoverable, not merely similar.

By Prashant C. Raju
Hugging Face Trending Papers
Aug 6

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at.

arXiv AI
Aug 20

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

The paper introduces the Graded Color Attribution (GCA) dataset, a benchmark that tests whether Vision‑Language Models (VLMs) and humans can articulate and follow a threshold rule for labeling objects by color. In experiments, humans consistently adhere to their stated rules, while VLMs—despite accurately estimating color coverage—often violate their own introspective rules, especially when world‑knowledge priors are present. This discrepancy highlights a miscalibration in VLM self‑knowledge that differs from human cognition.

By Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
arXiv Machine Learning
Jul 28

Beyond ICA: Identifiability by Symmetry Breaking

arXiv:2607. 23182v1 Announce Type: cross Abstract: We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting.

By Pengzhou Wu
arXiv Machine Learning
Sep 17

Transformation Laws in Neural Representations: Structure, Realisability, and Construction

The paper investigates how neural representations maintain the structure of input changes, linking representation analysis with internal interventions. It characterises when transformations can be applied through an encoder, providing linear settings where defects depend on discarded information and detailing failure modes for rectifiers and harmonic carriers. Using colour as a case study, the authors show that hue orbits in frozen visual features concentrate most energy in the first two harmonics, that this structure is inherited from input and architecture, and that a compact, fixed‑action interface can read hue zero‑shot with low error on unseen shapes.

By Yuan Sun
arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin