arXiv AI By Khawaja Murad ul Hassan, Mehran Ebrahimi

Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging

Read the original on arXiv AI →

The paper investigates why post‑hoc saliency maps, such as Grad‑CAM, shift when input images are rotated, even if the model’s prediction remains unchanged. By measuring equivariance at each stage of the CAM operator, the authors find that the instability originates from the spatial activation tensor rather than the channel weights, and that a training‑free wrapper called EquiGrad‑CAM can align and average saliency maps across rotated views to significantly improve rotation equivariance. Experiments on ImageNet, PatchCamelyon, and RESISC45 demonstrate that EquiGrad‑CAM outperforms rotation‑augmented training and enhances zero‑shot CLIP explanations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not

The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.

By Rub\'en Dar\'io Guerrero
Hugging Face Trending Papers
Jun 28

Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks

A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable.