arXiv Machine Learning

The Geometry of Updates: Fisher Alignment at Vocabulary Scale

arXiv:2606. 27242v1 Announce Type: new Abstract: Training-free source selection for LLM families with shared vocabularies arises in scientific string domains such as SMILES, protein, and genomic sequences, where candidate corpora share a tokenizer but differ in prediction targets.

arXiv AI
Sep 24

What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis

The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.

By Kentaro Oda
arXiv AI
Sep 24

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

The paper demonstrates that sharing a deep encoder alone does not eliminate the confounding effects in task-comparison scores. By introducing a conditional two‑discriminator discrepancy within the embedding space, the authors achieve robust detection of task changes, maintaining stability under input rotations and accurately tracking label‑permutation drift. This approach, integrated into a mixture‑of‑heads framework, outperforms traditional novelty triggers and generalizes across multiple backbones and datasets, including ImageNet‑21k ViT‑B/16, DINOv2, and CIFAR‑100.

By Kentaro Oda
arXiv Computation and Language
6d ago

Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders

The paper challenges the assumption that improving cross‑lingual alignment automatically enhances cross‑lingual transfer. Using XLM‑R models aligned on token, sentence, and masked‑language‑modeling objectives across four language pairs, the authors evaluate zero‑shot transfer on part‑of‑speech tagging and sentence classification. They find that embedding‑based alignment metrics poorly predict downstream performance and that alignment and task gradients are often nearly orthogonal, especially when operating at different representational levels.

By Yana Veitsman, Yihong Liu, Hinrich Sch\"utze
arXiv Machine Learning
Sep 25

The Alignment Illusion in Multimodal Large Language Models

The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.

By Hong-Han Wang, Yuntao Wang, Hu Ding
arXiv Machine Learning
Jun 18

Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment

arXiv:2606. 18703v1 Announce Type: new Abstract: Pretrained biological language models expose per-token probability distributions through masked-token prediction, providing the likelihood interface central to sequence design, variant scoring, and mechanistic interpretation.

By Yanjun Shao, Yundi Chen, Yashvi Patel, Aurelien Pelissier, Mar\'ia Rodr\'iguez Mart\'inez
arXiv Machine Learning
1d ago

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.

By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff