The paper presents a method for cross‑architecture knowledge distillation from a fine‑tuned DINOv2 Vision Transformer teacher to a lightweight bidirectional Visual State Space Model (LVSSM) student for tea leaf disease classification. By addressing training‑stability issues with a progressive convolutional stem and gated selective‑scan block, the 4.45 M‑parameter student achieves a mean test accuracy of 95.41%—a 3.09‑point improvement over the teacher’s 92.32%—while using only one‑fifth of the teacher’s parameters. Ablation studies show that simple logit‑level distillation outperforms intermediate feature alignment, and the gains are specific to students that start below the teacher’s performance.
By Zibo Zhou, Zongsen Qiu, Rui Chen, Yujie Yao, Yue Zhou, Jianjun Wang
The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.
By Amir Zamani, Zeinab Ghasemi-Naraghi
arXiv:2606. 14684v1 Announce Type: cross Abstract: Real-time fire classification systems require models that are simultaneously accurate, computationally efficient, and deployable on resource-constrained hardware.
By Mohammed Arif Mainuddin, Najifa Tabassum, Omar Ibne Shahid, Riasat Khan
arXiv:2609.23561v1 Announce Type: new
Abstract: Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical...
By Aleks Czufarow, Ihor Babin
The paper investigates Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) imagery, comparing modern convolutional neural networks (CNNs) and transformer-based deep neural networks (DNNs). It examines how factors such as network size, architecture, pretraining methods, data augmentation, and regularization influence performance, aiming to identify the highest-performing model and provide a training roadmap for state‑of‑the‑art SAS‑ATR systems.
By C. J. Moore, Alex Hurt, Jordan Malof
Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.
arXiv:2607. 16270v1 Announce Type: cross Abstract: Multilayer cloud detection from active--passive observation is vital for numerical weather prediction.
By Fu Wang, Chi Yang, Qi-Feng Lu, Rui-Xia Liu, Xiao-Fei Yang, Xiao-Fang Liu, Bo Li, Lin Chen
arXiv:2609.23061v1 Announce Type: new
Abstract: Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds....
By Linduo Wei, Junjie Fan, Yijun Mai, Yong Qi
arXiv:2609.06232v1 Announce Type: cross
Abstract: Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation. This paper repurposes them as spatial supervision...
By Sajjad Rezvani Boroujeni, Muskan Saraf, Gnana Tulasi Makineni, Tom Bush, Hossein Abedi
The paper presents a three‑stage, parameter‑efficient approach to improve automatic target recognition (ATR) with synthetic aperture sonar (SAS) data by adapting DINOv3 Vision Transformers. Stage 1 applies Low‑Rank Adaptation (LoRA) while freezing the backbone, which significantly boosts the area under the precision‑recall curve from 0.300 to 0.679. Subsequent hard‑negative mining and supervised contrastive learning stages show negligible impact, indicating that a single LoRA adaptation is sufficient for effective underwater ATR.
By Dan Zimmerman, Frank E. Bobe III, Amelia L. McCormack, Matthew Cook, Gregory D. Vetaw
arXiv:2609.13199v1 Announce Type: new
Abstract: Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher mo...
By Dawen Jiang, Zhishu Shen, Zeyu Liu, Tiehua Zhang
arXiv:2609.07915v1 Announce Type: cross
Abstract: Large vision models provide useful representations for remote-sensing segmentation but are often too expensive for deployment at the satellite or fie...
By Kishor Kumar Bhaumik, Nicolas Roque dos Santos, Jia Chen, Evangelos E. Papalexakis