DTKDP: A Dual Teacher Knowledge Distillation and Pruning Framework for Lightweight Oriented SAR Ship Detection
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper presents a method for cross‑architecture knowledge distillation from a fine‑tuned DINOv2 Vision Transformer teacher to a lightweight bidirectional Visual State Space Model (LVSSM) student for tea leaf disease classification. By addressing training‑stability issues with a progressive convolutional stem and gated selective‑scan block, the 4.45 M‑parameter student achieves a mean test accuracy of 95.41%—a 3.09‑point improvement over the teacher’s 92.32%—while using only one‑fifth of the teacher’s parameters. Ablation studies show that simple logit‑level distillation outperforms intermediate feature alignment, and the gains are specific to students that start below the teacher’s performance.
The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.
arXiv:2606. 14684v1 Announce Type: cross Abstract: Real-time fire classification systems require models that are simultaneously accurate, computationally efficient, and deployable on resource-constrained hardware.
arXiv:2609.23561v1 Announce Type: new Abstract: Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical...
The paper investigates Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) imagery, comparing modern convolutional neural networks (CNNs) and transformer-based deep neural networks (DNNs). It examines how factors such as network size, architecture, pretraining methods, data augmentation, and regularization influence performance, aiming to identify the highest-performing model and provide a training roadmap for state‑of‑the‑art SAS‑ATR systems.
Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.