arXiv AI By Yuxin Yue, Yingchen Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Xueqi Cheng

Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Aug 28

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

DEEPCHART is a new benchmark that evaluates large language models (LLMs) on faithful data‑science chart generation. It contains 1,482 expert‑annotated instances from scientific papers, financial filings, and ecosystem reports, and assesses chart creation through an Extract–Reason–Visualize pipeline. Experiments show that while LLMs can produce visually plausible charts, they frequently hallucinate data at the extraction and reasoning stages, especially in long, noisy, and multimodal contexts.

By Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
arXiv AI
Aug 10

Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning

arXiv:2608. 06938v1 Announce Type: cross Abstract: The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions.

By Chen Ling, Hanqian Li, Dongnan Liu, Keyu Qian, Jungang Li, Xinglong liu, Shiyi Wang, Xin Dong, Pengcheng Zhu, Wei Zhou, Linjian Mo, Nai Ding
arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
arXiv AI
Sep 12

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Mr.LHDR is a new benchmark designed to evaluate deep research agents on long‑horizon, multimodal tasks. It presents questions built from hidden Node‑Relation graphs that require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4 before arriving at a single verifiable answer. The benchmark tests both final answers and the correctness of intermediate conclusions, using metrics such as Overall Accuracy, Strict Accuracy, Checklist Score, and Dependency‑Aware Checklist Score.

By Minghao Guo, Meng Cao, Sui Zhao, Siyu Ning, Xin Wang, Haoze Zhao, Jiaxuan Yang, Haihong Hao, Mingfei Han, Shunlin Rong, Haijun Wu, Xiaodan Liang, Xiaojun Chang
arXiv Computer Vision
Aug 25

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA introduces a cache‑friendly, on‑demand visual evidence adaptation for multimodal speculative decoding, enabling a lightweight draft model to dynamically retrieve a bounded subset of visual memory based on a cumulative‑mass rule. The retrieved visual readout is fused with the draft hidden state via a lightweight gated residual correction, avoiding the insertion of visual tokens into the autoregressive context. Experiments on various vision‑language backbones and benchmarks show that FOVEA improves draft acceptance and speeds up end‑to‑end decoding by up to 2.13× compared to traditional autoregressive decoding.

By Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang