arXiv AI

Supervised Learning Has a Geometric Blind Spot

arXiv:2604. 21395v3 Announce Type: replace-cross Abstract: Ordinary supervised training minimises the task loss and then stops.

arXiv Machine Learning
Sep 23

Margin-Drop Coordinates for Cross-Budget Robustness Evaluation

arXiv:2609.26081v1 Announce Type: new Abstract: Fixed-budget robustness evaluation can select the wrong frozen vision encoder. An encoder that survives a shallow attack may lose most of that robustne...

By Yanliang Huang, Zhen Zhang, Peng Xie, Wenyuan Wu, Sitong Zhu, Zhuoqi Zeng, Amr Alanwar
arXiv AI
Sep 18

Information-Geometric Inverse Distillation for Enhancing Adversarial Transferability

The paper introduces Inverse Knowledge Distillation (IKD), an attack‑agnostic technique that enhances adversarial transferability by maximizing the discrepancy between benign and adversarial prediction distributions on a surrogate model. IKD employs a CE/KL‑equivalent soft‑label objective to push adversarial predictions away from a fixed benign anchor, leveraging Fisher‑sensitive surrogate directions. The authors provide theoretical analysis showing CE and KL induce identical gradients, derive a lower bound on Fisher‑subspace overlap, and demonstrate through extensive ImageNet experiments that IKD consistently improves black‑box attack performance across CNN, ViT, and defended models.

By Wenyuan Wu, Yuan Sun, Yingke Chen, Chao Su, Xi Peng, Dezhong Peng, Xu Wang
arXiv AI
Jun 18

Generalized Kullback-Leibler Divergence Loss

arXiv:2503. 08038v2 Announce Type: replace-cross Abstract: In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of (1) a weighted Mean Square Error (wMSE) loss and (2) a Cross-Entropy loss incorporating soft labels.

By Jiequan Cui, Beier Zhu, Qingshan Xu, Zhuotao Tian, Xiaojuan Qi, Bei Yu, Hanwang Zhang, Richang Hong
arXiv AI
6d ago

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.

By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv Machine Learning
Sep 21

On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation

The paper investigates the limits of the maximal coding rate reduction (MCR²) framework for out‑of‑distribution (OOD) generalisation. It shows that MCR² can lead to complete prediction failure under distribution shift, even when a perfectly stable feature is available, and that adding invariance principles from IRM or REx does not resolve this issue. The authors conclude that additional assumptions or learning principles are needed to guarantee stable OOD predictions with MCR².

By Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi, Po Yang