arXiv Computation and Language By Serwar Basch, Lizhen Qu, Iryna Gurevych

ReGround: Grounding Reviewer Comments in Multimodal Evidence

Read the original on arXiv Computation and Language →

ReGround is a new large‑scale dataset that links 10,267 reviewer comments to 16,274 pieces of evidence across 3,656 anonymous scientific papers, addressing the challenge of grounding comments in long multimodal documents. The dataset is constructed by leveraging explicit references in author rebuttals, providing high‑precision annotations. Evaluation shows that simple retrieval over full paper text performs poorly, evidence‑type inference is a major bottleneck, and multimodal evidence offers complementary signals that pure text misses.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 23

FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research

FMMD is a multimodal, multidisciplinary dataset of open peer reviews from F1000Research that pairs manuscript-level visual and structural data with version‑specific reviewer reports and editorial decisions. It addresses key gaps in existing datasets by preserving precise alignment between review comments and the exact manuscript version, and by including a wide range of scientific disciplines beyond computer science. The dataset supports tasks such as visual‑semantic consistency classification, figure‑related review comment generation, and editorial decision prediction, providing a comprehensive empirical resource for multimodal automated scholarly paper review research.

By Zhenzhen Zhuang, Yuqing Fu, Jing Zhu, Zhangping Zhou, Jialiang Lin
arXiv AI
Aug 5

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

arXiv:2608. 03292v1 Announce Type: new Abstract: Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages.

By Le Xiang, Zhicheng Guan, Hong Chen, Xiaocong Lin, Zhenghua Lei, Teng Hu, Bolei He, Long Zeng
arXiv AI
Sep 7

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

SciDocBench is a workflow-centered benchmark for scientific document understanding that includes 124 expert-authored questions across seven capability groups and 19 subtasks in five scientific domains. Each question is evaluated under four conditions—English or Chinese, all-images-first or interleaved document representations—resulting in 496 evaluation instances. The benchmark is paired with SciDocIR, a typed evidence-graph representation, and SciDocDataset, a collection of 15K fine-tuning and 8K reinforcement-learning samples, forming an evaluation-to-training framework for scientific-document assistants.

By Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, Dahua Lin
arXiv Computation and Language
Sep 18

Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.

By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu
arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou