arXiv:2603. 10834v3 Announce Type: replace-cross Abstract: Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes.
By Pum Jun Kim, Seung-Ah Lee, Seongho Park, Dongyoon Han, Jaejun Yoo
arXiv:2606. 03493v1 Announce Type: cross Abstract: Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets.
By Utku \c{S}irin, Cathy Hou, David Alvarez-Melis, Stratos Idreos
The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.
By Kirtan Rajesh
arXiv:2609.00868v1 Announce Type: cross
Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...
By Genpei Zhang
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid.
The paper investigates why vision‑language models like LLaVA‑1.5‑7B hallucinate objects in captions and proposes a targeted fix. By ranking attention heads whose image attention drops around hallucinated words, the authors identify 32 key heads and apply a head‑sliced LoRA adapter plus an inference‑time grounding controller. On COCO images, this combined method reduces hallucinated captions from 37% to 23% and hallucinated object mentions from 15.6% to 9.6%, while also lowering object recall.
By Armaan Sandhu, Abhilasha Senapati, Hima Kammachi