Learning Dynamic Evidence Routes for Vision Transformer Probing
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper audits whether routing entropy in Attention‑Residual transformer variants (Swin‑Tiny and DeiT‑Small) trained on CIFAR‑10/100 can signal prediction uncertainty beyond model confidence. Three tests examine the presence, consistency, and predictive power of routing signals, while a sensitivity audit measures how much injected effect the probes recover. Results show no significant improvement over confidence alone, with only modest recovery of injected signals and no consistent gains across seeds or metrics.
arXiv:2609.38956v1 Announce Type: new Abstract: Routing signals of modern vision transformers -- expert gates, attention-residual weights and halting scores -- often improve probes that predict wheth...
The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.
arXiv:2609.15579v1 Announce Type: new Abstract: An airborne sensor on a contested link cannot send video, so the appealing move is to send findings instead and report the ratio between the two. We ev...
arXiv:2609.38111v1 Announce Type: new Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when...
arXiv:2606. 08105v1 Announce Type: new Abstract: When attention concentrates on a single token, a sink, what is the model actually computing?