arXiv Machine Learning By Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou

Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling

Read the original on arXiv Machine Learning →

arXiv:2606. 26617v1 Announce Type: new Abstract: Scaling laws describe how learning performance varies with model size, data size, and compute.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 16

InfoNCE Induces Gaussian Distribution

arXiv:2602. 24012v2 Announce Type: replace Abstract: Contrastive learning has become a cornerstone of modern representation learning, allowing training with massive unlabeled data for both task-specific and general (foundation) models.

By Roy Betser, Eyal Gofer, Meir Yossef Levi, Guy Gilboa
arXiv Machine Learning
Aug 31

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.

By Mingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong