arXiv AI

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms

arXiv:2606. 30625v1 Announce Type: cross Abstract: Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes.

arXiv Machine Learning
Sep 1

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

The paper investigates why token prediction, a common pre‑training objective for language models, yields useful representations. It introduces a statistical framework linking token prediction accuracy to the geometry of token embeddings, showing that accurate predictions organize embeddings according to Hellinger distances between context distributions. The authors also propose a self‑consistency principle that refines contextual representations through repeated application of a shared block, and provide downstream guarantees for token generation, community recovery, and linear classification.

By Shulei Wang
arXiv Machine Learning
Jun 16

InfoNCE Induces Gaussian Distribution

arXiv:2602. 24012v2 Announce Type: replace Abstract: Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models.

By Roy Betser, Eyal Gofer, Meir Yossef Levi, Guy Gilboa
arXiv Machine Learning
Sep 16

Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation

The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.

By Micah Adler, John W. Byers, Mark Crovella
arXiv Machine Learning
Aug 31

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.

By Mingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong