The paper investigates few-shot fine-grained image classification and emphasizes the importance of phase information for capturing structural relationships. It introduces a plug‑and‑play amplitude‑phase integration (API) module that merges local and global frequency amplitude and phase data to create richer feature descriptors. A new network, PSF‑Net, adaptively fuses phase‑based spatial and frequency information and can be integrated into standard episodic training pipelines, achieving superior performance on five public datasets.
By Ruiling Liu, Linyue Zhang, Wenyi Zeng, Jiamiao Lu, Weichuang Zhang, Changming Sun, Zejun Zhang, Xiao Zhao
arXiv:2604. 16936v2 Announce Type: replace-cross Abstract: Feature reconstruction techniques are widely applied for few-shot fine-grained image classification (FSFGIC).
By Linyue Zhang, Wenyi Zeng, Zicheng Pan, Yongsheng Gao, Changming Sun, Jun Hu, Lixian Liu, Weichuan Zhang, Tuo Wang
arXiv:2610.01807v1 Announce Type: new
Abstract: Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existin...
By Ahmed Sharshar, Asif Hanif, Naveen Kumar Kummari, Mohammad Yaqub, Mohsen Guizan
arXiv:2607. 00251v1 Announce Type: cross Abstract: While most image deblurring techniques directly restore the spatial image variable, we propose an amplitude and phase decomposition recognizing the importance of accurate phase estimation in recovering sharp image details.
By Samira Malek, Haichuan Zhang, Chul Lee, Vishal Monga
The paper introduces ACF-Net, an optical flow‑guided framework for asymmetric audio‑visual fine‑grained visual categorization (FGVC), addressing challenges where video and audio are not strictly synchronized or matched. ACF-Net comprises Optical Flow‑Guided Motion (OFGM) to capture motion‑sensitive visual cues and suppress background noise, and Asymmetric Cross‑Modal Adaptive Fusion (ACAF) to estimate modality reliability and perform uncertainty‑aware fusion. The authors also present BirdPro, a new bird‑oriented audio‑visual benchmark with 1,919 audio recordings and 11,965 videos across 194 species, and report that ACF‑Net outperforms baselines by 2.97% in fused and 1.92% in mismatched settings.
By Bohan Deng, Shuo Ye, Zitong Yu
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha
The paper introduces LaST, a large-small collaborative framework for zero-shot surgical phase recognition. It combines a foundation model that generates frame-level phase priors with a lightweight model that refines predictions through iterative temporal refinement, dynamic quality control, and dual-model cross-learning. Experiments show LaST outperforms baseline and state-of-the-art methods, achieving significant accuracy gains on unseen clinical domains.
By Yiyi Zhang, Ying Zheng, Wenxin Fan, Yu Zhu, Yuchen Yuan, Litao Zhao, Zheng Li, Pheng-Ann Heng
The paper introduces Modality‑Specific Frequency Distillation (MSFD), a continual learning framework for video deepfake detection that separates spatial, temporal, and spatiotemporal features in the frequency domain. By preserving each modality independently and applying a cross‑modality decorrelation loss, MSFD adapts to new forgery patterns while maintaining performance across diverse continual deepfake video scenarios. Experiments demonstrate that this approach outperforms state‑of‑the‑art methods in both adaptation and retention.
By Taehoon Kim, Jongwook Choi, Heejae Jo, Byungmin Park, Jongwon Choi
arXiv:2605. 13672v1 Announce Type: cross Abstract: Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues.
By Giries Abu Ayoub, Morad Tukan, Loay Mualem
The paper introduces SGFNet, a Semantic‑Guided Fusion Network for classifying multi‑source remote sensing images. It features a Semantic Mixing Convolution Block that generates semantic‑aware kernels based on contextual relationships, and a Frequency Modulated Fusion Block that fuses cross‑modal information in the frequency domain to mitigate spatial misalignment. Experiments on the Augsburg and Houston 2018 datasets show SGFNet consistently outperforms state‑of‑the‑art methods.
By Yuwei Zhao, Chuanzheng Gong, Baogui Huan, Feng Gao, Junyu Dong, Qian Du
Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.
arXiv:2606. 23825v1 Announce Type: cross Abstract: Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operations of spatial-domain detectors that indiscriminately discard critical high-frequency details.
By Yuhan Rui, Shihan Qiao, Yibin Lou, Mingxi Yu, Yutong Wan, Yanqiao Chen, Dongsheng Hou, Zhen Cao, Athena Zhuoming Zhong, Qi Hao