arXiv Computer Vision

Spot-the-shift: Evaluating Grounded Image Difference Captioning of Long-term Changes

arXiv Machine Learning
Sep 4

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, allowing attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate these methods, especially for sparse, real‑world differences, with open‑weight models to ensure reproducibility.

By Julian Truetsch, Felix Hauser, Christoph Stiller, Frank Bieder
Hugging Face Trending Papers
Sep 3

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, enabling attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate this method, especially for sparse, real‑world differences, and the authors provide open‑weight models and code for reproducibility.

arXiv AI
Aug 17

EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning

arXiv:2608. 01856v2 Announce Type: replace Abstract: Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions.

By Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao
arXiv AI
Sep 4

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

The paper introduces Seeing Before Synthesizing (SBS), a weakly-supervised dense video captioning framework that uses a vision‑language model to generate frame‑level narratives for gaps between events and detect transitions based on semantic changes. SBS refines temporal masks by aligning transition points with vision‑language cues, rather than relying on rigidly placed synthetic captions. Experiments on ActivityNet Captions and YouCook2 show that SBS achieves state‑of‑the‑art results in both captioning and localization tasks.

By Ye-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim, Hwiseon Kim, Hyungee Kim, Dong-Jin Kim