arXiv AI By Chanho Park, Woochan Lee, Janyeong Oh, Geongho Gong, Minshu Kim, Yeachan Kwak, Seongim Choi

Early Cue Precision Shapes Visual Shortcut Learning in Controlled Cue-Manipulation Benchmarks

Read the original on arXiv AI →

arXiv:2606. 30344v1 Announce Type: cross Abstract: Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 24

TEEP-RCNN: Texture-Enhanced Edge-aware Perception for Steel Surface Defect Detection via Improved Convolutional Block Attention in Faster R-CNN

The paper introduces TEEP‑RCNN, a two‑stage detector that augments Faster R‑CNN with a Feature Pyramid Network backbone and an enhanced Convolutional Block Attention Module (CBAM) featuring dropout in the channel attention MLP and batch‑norm in the spatial attention branch. Training employs a differential learning‑rate schedule with cosine‑annealing warm‑up, and inference uses Test‑Time Augmentation combined with Weighted Box Fusion to stabilize localization of elongated and boundary‑adjacent defects. On the NEU‑DET benchmark, TEEP‑RCNN attains 73.3 % mAP@50 and 37.9 % mAP@50‑95 in only ten epochs on a single GPU, matching or surpassing YOLOv11m while excelling on the rolled‑in‑scale defect category under the COCO metric.

By Kirtan Rajesh
arXiv Computer Vision
Aug 27

Targeting the Attention Heads Behind Object Hallucination in LLaVA

The paper investigates why vision‑language models like LLaVA‑1.5‑7B hallucinate objects in captions and proposes a targeted fix. By ranking attention heads whose image attention drops around hallucinated words, the authors identify 32 key heads and apply a head‑sliced LoRA adapter plus an inference‑time grounding controller. On COCO images, this combined method reduces hallucinated captions from 37% to 23% and hallucinated object mentions from 15.6% to 9.6%, while also lowering object recall.

By Armaan Sandhu, Abhilasha Senapati, Hima Kammachi