Hugging Face Trending Papers

Learning a Flow to Self-Supervised Representations

The paper introduces Flow-Based Distribution Matching (FBDM), a non‑adversarial method that learns self‑supervised representations by aligning images to explicit geometric references through spherical conditional velocity regression. By using an ETF‑inspired reference, FBDM allows more reference components than the flow dimension while maintaining geometric separation, and it incorporates an alignment loss to bring augmented views closer together. Experiments on datasets from CIFAR to ImageNet show that FBDM performs nearly as well as adversarial distribution‑matching methods, achieves a 1.48‑ to 1.83‑fold speedup, and offers a theoretical bound on downstream misclassification rates.

arXiv Machine Learning
Sep 25

Learning a Flow to Self-Supervised Representations

The paper introduces Flow-Based Distribution Matching (FBDM), a non‑adversarial framework that learns self‑supervised representations using explicit geometric references and spherical conditional velocity regression. FBDM assigns augmented image views to shared target references while limiting reference usage, and employs an alignment loss to bring view representations closer. Experiments on datasets from CIFAR to ImageNet demonstrate that FBDM performs nearly as well as adversarial DM, outperforms existing SSL methods, and achieves a 1.48‑ to 1.83‑fold speedup with minimal GPU memory increase, while a theoretical analysis bounds downstream misclassification rates in terms of the pretraining loss.

By Yuling Jiao, Wensen Ma, Houduo Qi, Defeng Sun
arXiv Machine Learning
Aug 27

Continuous Adversarial Flow Models

The paper introduces continuous adversarial flow models, a continuous-time flow framework trained with an adversarial objective that replaces the fixed mean-squared-error criterion of flow matching. By incorporating a learned discriminator, the method guides training toward a different generalized distribution, yielding samples more closely aligned with the target data distribution. Applied as a post‑training step, it markedly improves ImageNet 256px generation metrics—reducing the guidance‑free FID of latent‑space SiT from 8.26 to 3.63 and of pixel‑space JiT from 7.17 to 3.57—and also enhances guided generation and text‑to‑image benchmarks.

By Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, Haoqi Fan
arXiv Computer Vision
Aug 31

Video Generative Models as Geometry Learner

The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.

By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
arXiv Computer Vision
5d ago

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources

The paper reports a controlled study of self‑supervised learning (SSL) objectives for image and video pretraining under limited data, architecture, and compute budgets. It compares contrastive, reconstruction, feature‑prediction, and diffusion methods, finding that DINOv2‑style pretraining delivers the best overall performance. Combining DINOv2 with video SSL objectives such as VideoMAE improves image classification and segmentation but harms video tracking and camera‑pose estimation, highlighting a trade‑off between semantic and geometric learning.

By Brun\'o B. Englert, Gijs Dubbelman
arXiv Machine Learning
Aug 27

JEPAMatch: Geometric Representation Shaping for Semi-Supervised Learning

JEPAMatch introduces a new semi‑supervised learning framework that replaces traditional output‑thresholding with explicit geometric shaping of latent representations. By combining the FlexMatch loss with a latent‑space regularization inspired by LeJEPA, the method encourages isotropic Gaussian structure in the embedding space, mitigating class imbalance and noisy pseudo‑labels. Experiments on CIFAR‑100, STL‑10, and Tiny‑ImageNet show consistent performance gains and faster convergence compared to existing FixMatch‑based baselines.

By Ali Aghababaei-Harandi, Aude Sportisse, Massih-Reza Amini
arXiv Computer Vision
Sep 16

MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking

MAETrack introduces a lightweight framework to adapt pretrained masked autoencoder (MAE) representations for 3D single object tracking (SOT). It uses Layer‑Selective Initialization (LSI) to keep shallow geometric layers from the pre‑training while re‑initializing deeper layers, and Geometric Residual Gating (GRG) to emphasize salient regions in BEV features before template‑search fusion. Experiments on standard 3D SOT benchmarks show consistent improvements over vanilla fine‑tuning with minimal computational cost.

By Sifan Zhou, Qiwei Wang, Linyue Tan, Ziyu Liu, Ziyu Zhao, Xiaobo Lu