Hugging Face Trending Papers

Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models

arXiv Machine Learning
Sep 23

Discovering Data Manifold Geometry through Geometric Properties

arXiv:2602.02611v2 Announce Type: replace Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...

By David Vigouroux (ANITI, IMT Atlantique - DSD, LaTIM), Lucas Drumetz (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Ronan Fablet (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Fran\c{c}ois Rousseau (IMT Atlantique - DSD, LaTIM)
arXiv Computer Vision
Sep 2

Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning

The paper introduces EquiSD, a label‑free training method that exploits scale equivariance to improve metric grounding in vision‑language models. By projecting model predictions onto a scale‑equivariant family and fine‑tuning on the resulting targets, EquiSD boosts a 3B model’s median response slope from 0.66 to 0.94 and raises mean relative accuracy by 9.2 points across simulated scales, with positive transfer to real QuantiPhy videos.

By Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu, Hanzhe Hong