arXiv Computer Vision
Sep 2

Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning

The paper introduces EquiSD, a label‑free training method that exploits scale equivariance to improve metric grounding in vision‑language models. By projecting model predictions onto a scale‑equivariant family and fine‑tuning on the resulting targets, EquiSD boosts a 3B model’s median response slope from 0.66 to 0.94 and raises mean relative accuracy by 9.2 points across simulated scales, with positive transfer to real QuantiPhy videos.

By Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu, Hanzhe Hong
arXiv AI
Jun 16

Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.

By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja
arXiv Computer Vision
Aug 27

Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding

The paper introduces Clue-OPSD, a clue‑privileged on‑policy self‑distillation framework that improves long‑video understanding by focusing on short, question‑relevant clue intervals rather than the entire video. Experiments on multiple benchmarks and Qwen3.5 model scales show that this approach consistently outperforms standard backbone models and competes strongly with supervised post‑training baselines, all while requiring fewer input frames and no additional inference modules.

By Kaishen Wang, Dongdi Zhao, Yijun Liang, Dingqiang Ye, Ruibo Chen, Heng Huang, Di Fu