On the Reliability of Cue Conflict and Beyond
arXiv:2603. 10834v3 Announce Type: replace-cross Abstract: Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes.
arXiv:2606. 30344v1 Announce Type: cross Abstract: Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression.
arXiv:2603. 10834v3 Announce Type: replace-cross Abstract: Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes.
arXiv:2606. 03493v1 Announce Type: cross Abstract: Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets.
The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.
arXiv:2609.00868v1 Announce Type: cross Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid.
The paper investigates why vision‑language models like LLaVA‑1.5‑7B hallucinate objects in captions and proposes a targeted fix. By ranking attention heads whose image attention drops around hallucinated words, the authors identify 32 key heads and apply a head‑sliced LoRA adapter plus an inference‑time grounding controller. On COCO images, this combined method reduces hallucinated captions from 37% to 23% and hallucinated object mentions from 15.6% to 9.6%, while also lowering object recall.
arXiv:2609.24565v2 Announce Type: replace Abstract: A connectome-constrained model of the fly visual system, optimized for motion and then frozen, can be driven over architectural drawings by prescri...
arXiv:2609.26168v1 Announce Type: cross Abstract: Recent work reports that vision--language models (VLMs) struggle to establish and maintain stable reference in repeated reference games. Rather than...
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
arXiv:2606. 01896v1 Announce Type: cross Abstract: Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expensive, or biased.
arXiv:2609.23655v1 Announce Type: new Abstract: Imbalanced gradient magnitudes between the visual and language pathways of a vision-language model are often treated as a defect to be corrected. We te...
arXiv:2606. 07882v1 Announce Type: cross Abstract: Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly different internal representations.