arXiv Computation and Language

ReGround: Grounding Reviewer Comments in Multimodal Evidence

ReGround is a new large‑scale dataset that links 10,267 reviewer comments to 16,274 pieces of evidence across 3,656 anonymous scientific papers, addressing the challenge of grounding comments in long multimodal documents. The dataset is constructed by leveraging explicit references in author rebuttals, providing high‑precision annotations. Evaluation shows that simple retrieval over full paper text performs poorly, evidence‑type inference is a major bottleneck, and multimodal evidence offers complementary signals that pure text misses.

arXiv Machine Learning
Sep 23

FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research

FMMD is a multimodal, multidisciplinary dataset of open peer reviews from F1000Research that pairs manuscript-level visual and structural data with version‑specific reviewer reports and editorial decisions. It addresses key gaps in existing datasets by preserving precise alignment between review comments and the exact manuscript version, and by including a wide range of scientific disciplines beyond computer science. The dataset supports tasks such as visual‑semantic consistency classification, figure‑related review comment generation, and editorial decision prediction, providing a comprehensive empirical resource for multimodal automated scholarly paper review research.

By Zhenzhen Zhuang, Yuqing Fu, Jing Zhu, Zhangping Zhou, Jialiang Lin
arXiv AI
Aug 5

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

arXiv:2608. 03292v1 Announce Type: new Abstract: Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages.

By Le Xiang, Zhicheng Guan, Hong Chen, Xiaocong Lin, Zhenghua Lei, Teng Hu, Bolei He, Long Zeng
arXiv AI
Sep 7

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

SciDocBench is a workflow-centered benchmark for scientific document understanding that includes 124 expert-authored questions across seven capability groups and 19 subtasks in five scientific domains. Each question is evaluated under four conditions—English or Chinese, all-images-first or interleaved document representations—resulting in 496 evaluation instances. The benchmark is paired with SciDocIR, a typed evidence-graph representation, and SciDocDataset, a collection of 15K fine-tuning and 8K reinforcement-learning samples, forming an evaluation-to-training framework for scientific-document assistants.

By Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, Dahua Lin
arXiv Computation and Language
Sep 18

Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.

By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu
arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
arXiv Computer Vision
Sep 18

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

DocAttriBench (DAB) is a large‑scale benchmark for fine‑grained, element‑level source attribution in Document Visual Question Answering (VQA). It introduces MAPPET, a Mask‑based Perplexity‑Derived Attribution method that uses document layout and language modeling to identify the most informative layout element for each answer. The benchmark contains 237k documents and 296k question‑answer pairs with element‑level grounding, and it evaluates multimodal LLMs on answer accuracy, attribution accuracy, and overall answer quality, revealing that even strong models often fail to localize supporting elements.

By Luca De Grandis (University of Modena and Reggio Emilia, Modena, Italy), Silvia Cappelletti (University of Modena and Reggio Emilia, Modena, Italy), William Raccagni (University of Modena and Reggio Emilia, Modena, Italy, University of Pisa, Pisa, Italy), Marcella Cornia (University of Modena and Reggio Emilia, Modena, Italy), Lorenzo Baraldi (University of Modena and Reggio Emilia, Modena, Italy), Rita Cucchiara (University of Modena and Reggio Emilia, Modena, Italy)
arXiv AI
Jun 16

VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA

arXiv:2606. 16092v1 Announce Type: cross Abstract: Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on multimodal large language models (MLLMs) for document QA predominantly produces text-only responses, underutilizing these visual elements.

By Young Rok Jang, Hyesoo Kong, Kyunghwan An, Jae Sub Huh, Gyeonghun Kim, Stanley Jungkyu Choi
arXiv Machine Learning
Sep 11

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

The paper introduces REVA, a method for compressing retrieval-augmented generation (RAG) prompts by aggregating historical query–document–model interactions into reusable evidence views. REVA mines attention traces from the target generator, maps token-level attention to readable words, aggregates importance across repeated document accesses, and produces budget‑specific plain‑text views that maintain document order and the standard RAG interface. Experiments on four benchmarks with modern LLMs show that REVA improves generation quality by 1.0–5.8 points over existing compressors while reducing compression overhead by 5.3 to 15.6 times and adding less than 40 ms of latency.

By Tuan Nguyen, Qiran Hu, Banruo Liu, Khoa D. Doan, Kok-Seng Wong, Fan Lai
arXiv Computation and Language
Sep 1

MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

MMDS-Bench is a new diagnostic benchmark for multimodal dynamic stance classification in social media parent‑reply interactions. It contains 3,482 multimodal instances annotated with a seven‑label stance taxonomy, plus an 800‑instance subset that demands structured reasoning over parent and reply understanding and stance‑relation inference. The benchmark also tags each instance with five challenge factors—multimodal fusion, parent framing, non‑literal expression, interaction reasoning, and label‑boundary ambiguity—and evaluates 12 multimodal large language models using a reference‑grounded LLM‑judge protocol, revealing that current models still struggle with relational inference beyond separate parent and reply comprehension.

By Yuzhe Ding, Kang He, Li Zheng, Shengwu Zheng, Teng Shi, Fei Li, Chong Teng, Donghong Ji