arXiv AI

Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging

The paper investigates why post‑hoc saliency maps, such as Grad‑CAM, shift when input images are rotated, even if the model’s prediction remains unchanged. By measuring equivariance at each stage of the CAM operator, the authors find that the instability originates from the spatial activation tensor rather than the channel weights, and that a training‑free wrapper called EquiGrad‑CAM can align and average saliency maps across rotated views to significantly improve rotation equivariance. Experiments on ImageNet, PatchCamelyon, and RESISC45 demonstrate that EquiGrad‑CAM outperforms rotation‑augmented training and enhances zero‑shot CLIP explanations.

arXiv Machine Learning
Sep 18

Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not

The paper introduces Stiefel Attention, which constrains the query and key projection matrices of transformers to the Stiefel manifold and optimizes them with a Riemannian Adam variant. It demonstrates that this approach yields steepest‑descent updates, is well‑conditioned, and preserves learned attention geometry during weight decay. Empirical results show significant accuracy gains on modular arithmetic grokking and CIFAR‑10 patches, with the improvement attributed to a step‑scale‑free update rule rather than equivariance or projector changes.

By Rub\'en Dar\'io Guerrero
Hugging Face Trending Papers
Jun 28

Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks

A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable.

arXiv Computer Vision
4d ago

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.

By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv Computer Vision
Sep 1

SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation

SNF-Bench is an evaluation framework for long‑horizon fixed‑camera video generation that separates static background fidelity from dynamic flow persistence and drift leakage. It reports these three factors independently, using controlled injections of translation, rotation, scale drift, and progressive freezing to validate each metric’s sensitivity. Auditing public checkpoints shows that whole‑frame motion metrics can mislead, while SNF‑Bench reveals the true trade‑offs between motion quality and background stability.

By Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park
arXiv Machine Learning
Jun 10

Post-Training Augmentation Invariance

arXiv:2505. 11702v3 Announce Type: replace Abstract: This work develops a framework for post-training augmentation invariance, in which our goal is to add invariance properties to a pretrained network without altering its behavior on the original, non-augmented input distribution.

By Keenan Eikenberry, Lizuo Liu, Yoonsang Lee