arXiv Computer Vision

Learning Dynamic Evidence Routes for Vision Transformer Probing

arXiv AI
1d ago

Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers

The paper audits whether routing entropy in Attention‑Residual transformer variants (Swin‑Tiny and DeiT‑Small) trained on CIFAR‑10/100 can signal prediction uncertainty beyond model confidence. Three tests examine the presence, consistency, and predictive power of routing signals, while a sensitivity audit measures how much injected effect the probes recover. Results show no significant improvement over confidence alone, with only modest recovery of injected signals and no consistent gains across seeds or metrics.

By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
arXiv Computation and Language
Sep 10

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

The paper introduces Contrastive Projection, a method that reads a transformer’s internal states by differencing the hidden states of two closely matched prompts and projecting the difference through the unembedding layer. This approach cancels shared components and highlights the distinctions between prompts, effectively revealing steering vectors and domain-to-domain mappings such as metaphor. The technique is training‑free, operates at every position, sub‑layer, and head, and has been validated across multiple architectures and initialization seeds.

By Olli Tuomi
arXiv AI
5d ago

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza
arXiv AI
Sep 12

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

ActMap is a new white‑box representation that compresses the entire hidden‑state trajectory of a language model during generation into a fixed 12 × 32 × 128 tensor. This compact 96 KiB map can be captured with no overhead and is read by a lightweight Vision Transformer to estimate answer correctness in a fraction of a millisecond. In experiments on short‑answer QA, math, and summarization, ActMap outperforms sampling, token‑probability, attention, and embedding baselines and matches a larger ACT‑ViT detector while achieving lower calibration error on most test pairs.

By Jacopo Dardini (University of Bologna), Roberta Calegari (University of Bologna)
Hugging Face Trending Papers
Sep 3

When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

The paper investigates how frozen vision‑language models (VLMs) respond to edits made directly to their internal object‑token representations, bypassing the image input. It introduces an answer‑key‑free protocol that evaluates edits by logical consistency and self‑audit, revealing that responses depend on explicit edit teaching, token cleanliness, density, and a separable reading axis. The study demonstrates that image‑free token edits can preserve most free‑text VQA performance and even outperform oracle methods on remote‑sensing benchmarks, with findings consistent across multiple datasets and model backbones.