arXiv:2609.03829v2 Announce Type: replace
Abstract: Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critica...
By Ruiling Liu, Linyue Zhang, Wenyi Zeng, Jiamiao Lu, Weichuang Zhang, Changming Sun, Zejun Zhang, Xiao Zhao
The paper investigates few-shot fine-grained image classification and emphasizes the importance of phase information for capturing structural relationships. It introduces a plug‑and‑play amplitude‑phase integration (API) module that merges local and global frequency amplitude and phase data to create richer feature descriptors. A new network, PSF‑Net, adaptively fuses phase‑based spatial and frequency information and can be integrated into standard episodic training pipelines, achieving superior performance on five public datasets.
By Ruiling Liu, Linyue Zhang, Wenyi Zeng, Jiamiao Lu, Weichuang Zhang, Changming Sun, Zejun Zhang, Xiao Zhao
arXiv:2607.05176v3 Announce Type: replace
Abstract: Small object detection (SOD) remains a challenging task in real-world applications. Despite recent advances, existing detectors remain limited by r...
By Aiwen Liu, Chengguang Zhu, Gang Wang, Dandan Zhu, Haodong Lin, Yan Wang, Huiyu Zhou, Zhengyi Pan
arXiv:2511. 10806v1 Announce Type: cross Abstract: Image deblurring is vital in computer vision, aiming to recover sharp images from blurry ones caused by motion or camera shake.
By Syed Mumtahin Mahmud, Mahdi Mohd Hossain Noki, Prothito Shovon Majumder, Abdul Mohaimen Al Radi, Md. Haider Ali, Md. Mosaddek Khan
arXiv:2606. 23825v1 Announce Type: cross Abstract: Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operations of spatial-domain detectors that indiscriminately discard critical high-frequency details.
By Yuhan Rui, Shihan Qiao, Yibin Lou, Mingxi Yu, Yutong Wan, Yanqiao Chen, Dongsheng Hou, Zhen Cao, Athena Zhuoming Zhong, Qi Hao
The paper introduces S$^3$F-Net, a dual‑branch network that fuses spatial and spectral representations for medical image classification. It combines a deep spatial CNN with a shallow spectral encoder, SpectraNet, which uses a learnable SpectralFilter layer to process the full Fourier spectrum efficiently. Evaluated on four medical imaging datasets, S$^3$F-Net consistently outperforms spatial‑only baselines, achieving state‑of‑the‑art accuracy on BRISC2025 and surpassing deeper models on the Chest X‑Ray Pneumonia dataset.
By Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
The paper introduces Modality‑Specific Frequency Distillation (MSFD), a continual learning framework for video deepfake detection that separates spatial, temporal, and spatiotemporal features in the frequency domain. By preserving each modality independently and applying a cross‑modality decorrelation loss, MSFD adapts to new forgery patterns while maintaining performance across diverse continual deepfake video scenarios. Experiments demonstrate that this approach outperforms state‑of‑the‑art methods in both adaptation and retention.
By Taehoon Kim, Jongwook Choi, Heejae Jo, Byungmin Park, Jongwon Choi
arXiv:2607. 17441v1 Announce Type: cross Abstract: Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust.
By Pamela Kirui, Cho Hyuk, Qingzhong Liu, Haodi Jiang
The paper introduces a Focal Log-Frequency Loss (f-loss) to counteract the spectral imbalance in pixel-space flow matching, where low frequencies dominate training. By balancing learning signals across frequencies and combining early frequency-domain supervision with later pixel-space refinement, the method accelerates convergence by up to 40% and improves FID and perceptual fidelity across multiple model scales. It requires no architectural changes and can replace existing flow matching losses as a drop‑in solution.
By Lucas Degeorge, Paul Couairon, Arijit Ghosh, Alexei A. Efros, David Picard, Vicky Kalogeiton
The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significant modal disparities caused by distinct imaging mechanisms. These intrinsic inconsistencies are prone to introducing pseudo-changes, thereby constraining detection accuracy.
arXiv:2606. 08204v1 Announce Type: new Abstract: Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities.
By Alonso Urbano, David W. Romero, Max Zimmer, Sebastian Pokutta
arXiv:2606. 28226v1 Announce Type: cross Abstract: Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepancies between training and inference.
By Guanbo Huang, Jingjia Mao, Fanding Huang, Fengkai Liu, Xiangyang Luo, Yaoyuan Liang, Jiasheng Lu, Xiaoe Wang, Pei Liu, Ruiliu Fu, Ruqi Huang, Shao-Lun Huang