arXiv Computer Vision By Anthony Fuller, Scott C. Lowe, Daniel G. Kyrollos, Graham W. Taylor, Evan Shelhamer, James R. Green

Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

arXiv Computer Vision
1d ago

Scalable Patch-Level Self-Supervised Learning

arXiv:2610.10013v1 Announce Type: new Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of mul...

By Maximilian Seitzer, Gabriele Trivigno, Anton\'in Vobeck\'y, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Sim\'eoni, Piotr Bojanowski
arXiv AI
2d ago

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.

By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie