arXiv Computation and Language

MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

MMDS-Bench is a new diagnostic benchmark for multimodal dynamic stance classification in social media parent‑reply interactions. It contains 3,482 multimodal instances annotated with a seven‑label stance taxonomy, plus an 800‑instance subset that demands structured reasoning over parent and reply understanding and stance‑relation inference. The benchmark also tags each instance with five challenge factors—multimodal fusion, parent framing, non‑literal expression, interaction reasoning, and label‑boundary ambiguity—and evaluates 12 multimodal large language models using a reference‑grounded LLM‑judge protocol, revealing that current models still struggle with relational inference beyond separate parent and reply comprehension.

arXiv Computation and Language
4d ago

Embedding Models for Stance-Aware Argument Retrieval

The paper investigates how dense embedding models can be used for stance-aware argument retrieval, a task that requires both topic relevance and correct stance (support or attack) toward a claim. Experiments reveal that current models favor topical overlap and ignore stance, and that contrastive training to fix this bias leads to over-correction, where models focus too much on polarity keywords at the expense of topic relevance. To address this, the authors propose diagnostic word-ablation metrics and a data‑centric solution involving a balanced argument curriculum and LLM‑augmented stance‑inverted arguments, which helps powerful models learn deeper directional logic and improves stance‑aware retrieval performance.

By Angelo Sparacino, Francesca Toni, Adam Dejl
arXiv AI
Jun 10

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

arXiv:2606. 10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models.

By Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad, Ufaq Khan, Muhammad Haris Khan
arXiv AI
2d ago

Visual Framing for News Stance Detection via Image Generation

The paper introduces VFStance, a method that uses image generation to make implicit stance cues in news articles more explicit through visual framing. It targets article-level news stance detection, a task complicated by subtle, structurally complex texts. Experiments show VFStance outperforms existing methods, and a user study with 200 participants demonstrates that the visual framing makes stance signals more noticeable in a snippet-based news consumption setting.

By Dahyun Lee, Jiyoung Han, Kunwoo Park
arXiv AI
1d ago

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

DocHop is a new benchmark that tests multimodal large language models on integrated chart‑context reasoning within document‑style images. The benchmark presents narrative text that imposes multi‑step compositional constraints, while charts supply the data needed to answer questions grounded in semantic reference labels. It contains 2,074 examples across six task categories, generated via a stochastic logic‑first pipeline that controls reasoning depth and visual density, and shows a large performance gap between humans (over 90% accuracy) and the best models (62.83%).

By Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee
arXiv AI
Jun 9

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

arXiv:2606. 09169v1 Announce Type: new Abstract: In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.

By Lingyi Meng, Zecong Tang, Haoran Li, Tengju Ru, Zhejun Cui, Weitong Lian, Qi Kang, Hangshuo Cao, Yichen Zhu, Yechi Liu, Kaixuan Wang, Yu-Jie Yuan, Chunwei Wang, Yu Zhang, Bo Dai