arXiv AI

CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications

arXiv:2608. 15972v1 Announce Type: cross Abstract: Synchronized camera and wireless measurements observe the same scene through different physical channels.

arXiv Computer Vision
Aug 28

Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification

The paper presents a method for cross‑architecture knowledge distillation from a fine‑tuned DINOv2 Vision Transformer teacher to a lightweight bidirectional Visual State Space Model (LVSSM) student for tea leaf disease classification. By addressing training‑stability issues with a progressive convolutional stem and gated selective‑scan block, the 4.45 M‑parameter student achieves a mean test accuracy of 95.41%—a 3.09‑point improvement over the teacher’s 92.32%—while using only one‑fifth of the teacher’s parameters. Ablation studies show that simple logit‑level distillation outperforms intermediate feature alignment, and the gains are specific to students that start below the teacher’s performance.

By Zibo Zhou, Zongsen Qiu, Rui Chen, Yujie Yao, Yue Zhou, Jianjun Wang
arXiv AI
Jul 14

JEPA for AI-Native 6G: Predictive Representations and Open Challenges

arXiv:2607. 09798v1 Announce Type: cross Abstract: Sixth-generation (6G) networks are moving toward AI-native operation, where learning modules are embedded across the radio access network (RAN), edge, and core.

By Sheikh Salman Hassan, Irshad A. Meer, Almoatssimbillah Saifaldawla, Yan Kyaw Tun, Mustafa Ozger, Madyan Alsenwi, Nguyen Van Huynh, Woong-Hee Lee, Cedomir Stefanovic, Mathini Sellathurai, Henk Wymeersch, Tharmalingam Ratnarajah
arXiv Computer Vision
6d ago

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

arXiv:2609.31620v1 Announce Type: new Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual r...

By Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
arXiv Computer Vision
Sep 3

SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging

SlowFast‑SCI introduces a dual‑speed deep‑unfolding framework for spectral compressive imaging that combines a slow, pre‑trained backbone with a fast, test‑time adaptation stage. The slow phase distills a priors‑based model into a compact fast‑unfolding network, while the fast phase embeds lightweight modules that self‑supervise at test time without retraining the backbone. This design yields significant reductions in parameters and FLOPs, improves out‑of‑distribution PSNR by up to 5.79 dB, and accelerates adaptation four‑fold, all while remaining modular enough to integrate with any existing deep‑unfolding system.

By Haijin Zeng, Xuan Lu, Jiezhang Cao, Kai Zhang, Yurong Zhang, Qiangqiang Shen, Guoqing Chao, Li Jiang, Yongyong Chen, Jingyong Su, Jie Liu
arXiv Computer Vision
Sep 17

STRADAViT: Self-Supervised Domain Adaptation of Vision Transformer Backbones for Radio Astronomy

STRADAViT is a self‑supervised continued‑pretraining framework that adapts Vision Transformer (ViT) backbones for radio‑astronomy image analysis. It curates mixed‑survey data, generates radio‑astronomy‑aware training views, and initializes encoders with ViT‑MAE, optionally adding register tokens. Evaluations on three morphology benchmarks (MiraBest, LoTSS DR2, and Radio Galaxy Zoo) show that a register‑based two‑stage checkpoint improves linear‑probe Macro‑F1 scores over the ViT‑MAE baseline and enhances fine‑tuning on MiraBest and RGZ DR1, though performance on LoTSS DR2 fine‑tuning declines; these differences are statistically significant.

By Andrea DeMarco, Ian Fenech Conti, Hayley Camilleri, Ardiana Bushi, Simone Riggi
arXiv AI
3d ago

Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking

The paper introduces TSDA-Track, a Template-Search Domain Adaptation framework designed to reduce modality gaps in cross‑modal visual object tracking. Two variants are explored: Pre‑AFA TSDA‑Track uses adversarial alignment before transformer interaction, while Enc‑CFA TSDA‑Track applies contrastive alignment after interaction to strengthen cross‑modal correspondence. Experiments on datasets such as LasHeR, RGBT234, GTOT, and Anti‑UAV‑024 show that both variants outperform state‑of‑the‑art trackers, with Pre‑AFA achieving an SR/PR of 43.2/56.0 on RGBT234 under the modality‑switch protocol.

By Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran
arXiv Machine Learning
Jul 14

TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation

arXiv:2607. 10762v1 Announce Type: cross Abstract: Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged as a self-supervised pretraining strategy that reduces reliance on dense point-wise annotation for 3D scene understanding.

By Sutharsan Mahendran, Darshana Priyasad, Kaushik Roy, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam