The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.
By Wenyuan Wu, Yuan Sun, Yingke Chen, Chao Su, Xi Peng, Dezhong Peng, Xu Wang
arXiv:2608.30699v1 Announce Type: cross
Abstract: Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to...
By Yue Cheng, Jiajun Zhang, Xiaohui Gao, Weiwei Xing, Zhanxing Zhu
The paper investigates how the sampled-token reverse-KL loss in on‑policy distillation distributes updates across tokens. By analyzing the gradient of the per‑token K2 estimator, the authors find that tokens with low student probability and large teacher‑student gaps receive disproportionately large gradient norms. They propose Surprise‑aware Reweighting (SuRe), a lightweight weighting rule that further amplifies this allocation, and demonstrate that SuRe improves math metrics on Qwen3 student models without harming out‑of‑domain performance.
By Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo, Yuming Yang, Mingxu Chai, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang
The paper introduces DUA-D2C, a Dynamic Uncertainty-Aware Divide2Conquer method that improves overfitting remediation in deep learning. It refines the traditional Divide2Conquer approach by dynamically weighting subset models based on a composite score of accuracy and normalized prediction entropy, allowing the central model to learn more from generalizable and confident edge models. The authors provide theoretical justification, show reduced model variance, and demonstrate significant generalization gains across image, audio, and text benchmarks, even when combined with standard regularizers like Dropout.
By Md. Saiful Bari Siddiqui, Md Mohaiminul Islam, Md. Golam Rabiul Alam
arXiv:2606. 16050v1 Announce Type: cross Abstract: Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE) exhibit unbounded sensitivity to outliers.
By Mainak Kundu, Ria Kanjilal, Ismail Uysal
The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.
By Kabir Thayani
arXiv:2606. 31664v1 Announce Type: cross Abstract: Performance in face and speaker verification is largely driven by margin-penalty softmax losses such as CosFace and ArcFace.
By Dimitrios Koutsianos, Ladislav Mo\v{s}ner, Yannis Panagakis, Themos Stafylakis
arXiv:2608. 15215v1 Announce Type: cross Abstract: Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass.
By Gordei Verbii, Juho Lee
The paper studies weak-to-strong generalization (W2SG), where a student model trained on a weaker teacher’s labels surpasses the teacher on the target task. Using a Bregman divergence bias‑variance decomposition, it shows that the student‑teacher risk gap depends on their expected misfit, without requiring convexity of the student hypothesis class. For squared loss, a sufficient condition is that the student converges to the teacher’s posterior mean, achievable by enlarging the student; for cross‑entropy loss, reducing the student’s predictive entropy and using reverse cross‑entropy can promote W2SG, which is empirically validated.
By Gengze Xu, Wei Yao, Ziqiao Wang, Yong Liu
arXiv:2505. 03201v4 Announce Type: replace-cross Abstract: Integrated Gradients (IG) is a widely used attribution method in explainable AI, particularly in computer vision applications where reliable feature attribution is essential.
By Kien Tran Duc Tuan, Tam Nguyen Trong, Son Nguyen Hoang, Khoat Than, Anh Nguyen Duc
VICAL is a framework for long‑tailed visual recognition that focuses on reducing prediction variance rather than increasing expert diversity. It combines Self‑Consistency Learning, which smooths the loss landscape and mitigates overfitting on tail classes, with Deep Ensemble Distillation, which encourages low‑frequency semantic agreement across experts. Experiments on CIFAR‑LT, ImageNet‑LT, and iNaturalist 2018 demonstrate that VICAL consistently outperforms state‑of‑the‑art methods.
By Jiangang Zhu, Zheng Wang, Bin Zhu, Yi-Ping Phoebe Chen, Jingjing Chen
arXiv:2606. 01746v1 Announce Type: cross Abstract: Modern neural networks are highly susceptible to adversarial perturbations.
By Kai Wang