arXiv Computation and Language

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

The paper introduces Data Journalist Agent (Data2Story), a multi‑agent framework that orchestrates specialized roles into a single virtual newsroom to produce evidence‑grounded, multimodal news stories. It ensures every claim is traceable to data, code, or external references via an Inspector, and generates interactive visualizations such as maps and audio to match reader interests. Evaluations on 18 articles show competitive performance in angle coverage, rubric scores, and verifiability, while human writers still lead in editorial angle and creative design.

arXiv AI
Jun 19

DataMagic: Transforming Tabular Data into Data Insight Video

arXiv:2606. 20388v1 Announce Type: cross Abstract: Data videos integrate dynamic charts, voice narration, and synchronized animations to communicate data insights as temporal narratives, making them an effective medium for improving data consumption efficiency in the data management lifecycle.

By Yupeng Xie, Chen Ma, Zhenyang Wang, Liangwei Wang, Jiayi Zhu, Chuxuan Zeng, Zhouan Shen, Boyan Li, Yuyu Luo
arXiv AI
Aug 5

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

arXiv:2608. 03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia.

By Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
arXiv AI
Jul 7

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

arXiv:2607. 04438v1 Announce Type: cross Abstract: Research dissemination, turning a paper into a poster, a talk video, and a blog post, is still a manual last mile.

By Lingao Xiao, Yalun Dai, Yangyu Huang, Qihao Zhao, Wenshan Wu, Hugo He, Ruishuo Chen, Jin Jiang, Qianli Ma, Jiahuan Zhang, Xin Zhang, Ying Xin, Yang Ou, Yan Xia, Scarlett Li, Longbo Huang, Zhipeng Zhang, Yang He, Yap Kim Hui, Yan Lu
arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
arXiv AI
Sep 1

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

The paper introduces Agent-as-Peer-Debriefing, a multi‑agent framework that incorporates peer debriefing into qualitative data analysis with large language models. A Hierarchical Coding Agent generates codes and reflections, which are then refined by three Peer‑Debriefing Agents applying Theory‑Driven, Data‑Driven, or Applied perspectives. Experiments on three datasets show that perspective‑based refinement aligns more closely with human codes than a single‑LLM baseline, and that the choice of perspective offers meaningful trade‑offs.

By Zhimin Lin, Kun Cheng, Zhiyao Shu, Junhua Fang, Juntao Li, Fan Bai, Jie Gao
Hugging Face Trending Papers
Jul 5

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

Research dissemination, turning a paper into a poster, a talk video, and a blog post, is still a manual last mile. Prior automation treats each artifact in isolation that each re-extract the paper from scratch, usually ship one-way renders the author cannot reopen in PowerPoint or Word, and gates quality on soft VLM-preference scores that plateau while load-bearing sections still read as empty.

arXiv Computation and Language
6d ago

Epstein Files Engine: Agentic Search for Investigative Journalism

The Epstein Files Engine is an AI agent developed by the New York Times to help journalists investigate a massive mixed‑media collection released by the U.S. Department of Justice on January 30, 2026, which contains about three million pages of PDFs related to Jeffrey Epstein. The Engine translates reporter questions into Google BigQuery SQL queries across three corpora—Epstein‑related releases, the Times’s archive, and external Epstein‑related news headlines—using an LLM to plan queries and return citation‑rich answers that reporters can verify. Over 100 journalists used the Engine, contributing to at least 20 published stories, and the system includes a Diff method for text‑and‑visual duplicate matching to surface genuinely new information.

By Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward