Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA). While recent approaches predominantly rel...
arXiv:2609.17458v1 Announce Type: cross
Abstract: Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering...
By Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi, Utkarsh Verma, Ganesh Ramakrishnan
The paper investigates pixel-level table compression for document question answering, comparing five vision‑language models across two benchmarks and varying visual‑token budgets. It finds that representing tables as native‑resolution images matches text in performance and efficiency, while highly downscaled images still allow the model to identify relevant tables but lose readability, leading to longer reasoning traces. A two‑step, training‑free method first selects relevant tables from compressed images and then reasons over them at native resolution, saving 41% of tokens and improving accuracy by 7 points over single‑step native‑resolution QA, while using 15% fewer tokens than the most efficient single‑step compressed setup without accuracy loss.
By I\~nigo Alonso, Mirella Lapata
arXiv:2601. 04498v2 Announce Type: replace Abstract: Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information.
By Yinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen
arXiv:2609.13158v1 Announce Type: new
Abstract: Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar...
By Yongqi Yu, Yu Zhang
The paper introduces TEmBed, a unified benchmark for evaluating tabular embeddings across four representation levels—cell, row, column, and table—using a diverse set of models. It demonstrates that the best model depends on the specific task and representation level, providing practical guidance for selecting embeddings in real-world applications. The study aims to facilitate the development of more general-purpose tabular representation models.
By Liane Vogel, Kavitha Srinivas, Niharika D'Souza, Sola Shirai, Oktie Hassanzadeh, Horst Samulowitz