The paper introduces Flow-Based Distribution Matching (FBDM), a non‑adversarial method that learns self‑supervised representations by aligning images to explicit geometric references through spherical conditional velocity regression. By using an ETF‑inspired reference, FBDM allows more reference components than the flow dimension while maintaining geometric separation, and it incorporates an alignment loss to bring augmented views closer together. Experiments on datasets from CIFAR to ImageNet show that FBDM performs nearly as well as adversarial distribution‑matching methods, achieves a 1.48‑ to 1.83‑fold speedup, and offers a theoretical bound on downstream misclassification rates.
arXiv:2502. 14424v3 Announce Type: replace-cross Abstract: Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified.
By Yuling Jiao, Wensen Ma, Defeng Sun, Hansheng Wang, Yang Wang
The paper introduces continuous adversarial flow models, a continuous-time flow framework trained with an adversarial objective that replaces the fixed mean-squared-error criterion of flow matching. By incorporating a learned discriminator, the method guides training toward a different generalized distribution, yielding samples more closely aligned with the target data distribution. Applied as a post‑training step, it markedly improves ImageNet 256px generation metrics—reducing the guidance‑free FID of latent‑space SiT from 8.26 to 3.63 and of pixel‑space JiT from 7.17 to 3.57—and also enhances guided generation and text‑to‑image benchmarks.
By Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, Haoqi Fan
arXiv:2602.05391v3 Announce Type: replace
Abstract: Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for do...
By Qianxin Xia, Jiawei Du, Yuhan Zhang, Xin Zhang, Xuewan He, Wenbo Jiang, Jielei Wang, Tao Luo, Guoming Lu
The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.
By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
arXiv:2605.06272v2 Announce Type: replace
Abstract: While generative modeling has achieved remarkable success on tasks like natural language-conditioned image generation, enabling model adaptation fr...
By Tyler Ingebrand, Ruihan Zhao, Kushagra Gupta, David Fridovich-Keil, Sandeep P. Chinchali, Ufuk Topcu
arXiv:2609.36348v1 Announce Type: cross
Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
By Xiaoyu Wu, Yifei Wang, Chen Wei
arXiv:2602. 23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
By Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata
The paper reports a controlled study of self‑supervised learning (SSL) objectives for image and video pretraining under limited data, architecture, and compute budgets. It compares contrastive, reconstruction, feature‑prediction, and diffusion methods, finding that DINOv2‑style pretraining delivers the best overall performance. Combining DINOv2 with video SSL objectives such as VideoMAE improves image classification and segmentation but harms video tracking and camera‑pose estimation, highlighting a trade‑off between semantic and geometric learning.
By Brun\'o B. Englert, Gijs Dubbelman
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
arXiv:2609.35763v3 Announce Type: replace
Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representa...
By Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
arXiv:2606. 10089v1 Announce Type: cross Abstract: In this work, we develop theoretical foundation for flow matching with neural-network-parameterized conditional velocity fields.
By Yihan He, Qishuo Yin, Yuan Cao, Jianqing Fan, Han Liu