arXiv:2609.03829v2 Announce Type: replace
Abstract: Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critica...
By Ruiling Liu, Linyue Zhang, Wenyi Zeng, Jiamiao Lu, Weichuang Zhang, Changming Sun, Zejun Zhang, Xiao Zhao
arXiv:2604. 16936v2 Announce Type: replace-cross Abstract: Feature reconstruction techniques are widely applied for few-shot fine-grained image classification (FSFGIC).
By Linyue Zhang, Wenyi Zeng, Zicheng Pan, Yongsheng Gao, Changming Sun, Jun Hu, Lixian Liu, Weichuan Zhang, Tuo Wang
arXiv:2610.01807v1 Announce Type: new
Abstract: Reliable clinical deployment of deep medical image models is hindered by distribution shifts across scanners, sites, and acquisition protocols. Existin...
By Ahmed Sharshar, Asif Hanif, Naveen Kumar Kummari, Mohammad Yaqub, Mohsen Guizan
arXiv:2607. 00251v1 Announce Type: cross Abstract: While most image deblurring techniques directly restore the spatial image variable, we propose an amplitude and phase decomposition recognizing the importance of accurate phase estimation in recovering sharp image details.
By Samira Malek, Haichuan Zhang, Chul Lee, Vishal Monga
The paper introduces LaST, a large-small collaborative framework for zero-shot surgical phase recognition. It combines a foundation model that generates frame-level phase priors with a lightweight model that refines predictions through iterative temporal refinement, dynamic quality control, and dual-model cross-learning. Experiments show LaST outperforms baseline and state-of-the-art methods, achieving significant accuracy gains on unseen clinical domains.
By Yiyi Zhang, Ying Zheng, Wenxin Fan, Yu Zhu, Yuchen Yuan, Litao Zhao, Zheng Li, Pheng-Ann Heng
The paper introduces ACF-Net, an optical flow‑guided framework for asymmetric audio‑visual fine‑grained visual categorization (FGVC), addressing challenges where video and audio are not strictly synchronized or matched. ACF-Net comprises Optical Flow‑Guided Motion (OFGM) to capture motion‑sensitive visual cues and suppress background noise, and Asymmetric Cross‑Modal Adaptive Fusion (ACAF) to estimate modality reliability and perform uncertainty‑aware fusion. The authors also present BirdPro, a new bird‑oriented audio‑visual benchmark with 1,919 audio recordings and 11,965 videos across 194 species, and report that ACF‑Net outperforms baselines by 2.97% in fused and 1.92% in mismatched settings.
By Bohan Deng, Shuo Ye, Zitong Yu
arXiv:2603. 21378v2 Announce Type: replace-cross Abstract: Phase unwrapping remains a critical and challenging problem in InSAR processing, particularly in scenarios involving complex deformation patterns.
By Yijia Song, Juliet Biggs, Alin Achim, Robert Popescu, Simon Orrego, Nantheera Anantrasirichai
arXiv:2605. 13672v1 Announce Type: cross Abstract: Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues.
By Giries Abu Ayoub, Morad Tukan, Loay Mualem
arXiv:2504.10079v5 Announce Type: replace
Abstract: Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level repre...
By Hongyu Qu, Ling Xing, Jiachao Zhang, Rui Yan, Yazhou Yao, Xiangbo Shu
Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.
arXiv:2608. 08963v1 Announce Type: cross Abstract: Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data.
By Sarah Rastegar, Mina Ghadimi Atigh, Pascal Mettes, Yuki M. Asano, Cees G. M. Snoek
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha