arXiv:2601.13111v3 Announce Type: replace-cross
Abstract: Realistic text-to-SQL workflows often require joining multiple tables. As a result, accurately retrieving the relevant set of tables becomes...
By Hassan Soliman, Vivek Gupta, Dan Roth, Iryna Gurevych
Tables are ubiquitous across diverse domains, yet reasoning over them remains a significant challenge for modern large language models (LLMs). Current approaches typically linearize tables into sequen...
H2Table introduces a hierarchical hypergraph representation for complex tables, enabling a hypergraph encoder to capture semantic relationships between headers and cells. The framework uses learnable query vectors to extract structural embeddings for large language models. Experiments on the HiTab dataset show a 22.88% improvement over state‑of‑the‑art baselines on tables with four levels of nesting.
By Jia Ling, Yangfan Wang, Chen Tang, Haoming Tan, Yang Yang, Yi Guan, Jingchi Jiang
arXiv:2610.00817v1 Announce Type: cross
Abstract: Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstrea...
By Sandipan De, Jin Wang, Vivek Gupta
arXiv:2607. 06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.
By So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen
PARTAB is a framework that improves large language model reasoning on tables by constructing a structured evidence interface. It represents query‑relevant evidence as semantically coherent, row‑linked table regions and performs hierarchical selection over column groups and row‑level partitions before composing the evidence for answer generation. Evaluations on multiple table reasoning benchmarks show that PARTAB consistently outperforms full‑table prompting and recent methods, achieving strong performance on WikiTableQuestions and TabFact while remaining competitive on numerical reasoning tasks.
By Md Mahadi Hasan Nahid, Davood Rafiei
MetaRTL is a two-stage framework for relational table learning that first generates lightweight pre-trained table embeddings and then applies non‑parametric message passing to extract meta‑path features. These features are aggregated using an attention module called MetaAttn, shifting computation from deep GNN stacks to efficient meta‑path aggregation. Experiments on 10 real‑world datasets across 24 tasks show that MetaRTL achieves strong performance while reducing computational cost.
By Ken Zhong, Weichen Li, Zheng Wang
arXiv:2411. 19504v2 Announce Type: replace Abstract: The advance of large language models (LLMs) has unlocked great opportunities in complex multi-modal data management tasks, particularly in question answering (QA) over complicated multi-table relational data.
By Zipeng Qiu, Chenyue Li, You Peng, Guangxin He, Binhang Yuan, Chen Wang
arXiv:2608. 03565v1 Announce Type: new Abstract: While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoring the rich semantics inherent in feature names or cell entries.
By G\"unther Schindler, Maximilian Schambach, Johannes H\"ohne
InRTL: Effective Intra-Inter Interaction Learning for Relational Tables proposes a unified framework that explicitly models dependencies both within and across relational tables. The approach introduces intra-table interactions to capture associations among rows in the same table and inter-table interactions to capture dependencies between rows linked by primary key–foreign key relationships. It employs a column-aware table encoder, Transformer-based self-attention and cross-attention modules, and incorporates linearized attention and heterogeneous graph neural networks to improve scalability, demonstrating effectiveness across ten datasets and 24 real-world tasks.
By Weichen Li, Ken Zhong, Zheng Wang, Li Pan, Jianhua Li
arXiv:2606. 28916v1 Announce Type: cross Abstract: We introduce GRAB, a constructor-encoder-bridge pipeline for table question answering.
By Simone Varriale, Tamara Cucumides, Floris Geerts, Paolo Papotti
The paper introduces FlockMTL, an extension for database management systems that deeply integrates large language models and retrieval‑augmented generation into DuckDB. It provides model‑driven scalar and aggregate functions, cost‑based optimizations like batching and caching, and new SQL DDL abstractions (PROMPT and MODEL) to treat LLMs as first‑class schema objects. These features aim to simplify the development of knowledge‑intensive analytical applications by reducing the effort required to orchestrate heterogeneous data systems and manage LLM context.
By Anas Dorbani, Sunny Yasser, Jimmy Lin, Amine Mhedhbi