arXiv Machine Learning

Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation

arXiv:2606. 01890v1 Announce Type: new Abstract: Real-world domains often contain heterogeneous tables whose headers vary while their underlying attribute semantics are shared, making it difficult to induce domain-specialized semantics from table-local evidence alone.

arXiv AI
Jun 9

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

arXiv:2606. 09323v1 Announce Type: new Abstract: Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare directly even when they operate on similar tabular signals.

By Wei Pang, Xiangru Jian, Hehan Li, Zhixuan Yu, Alex Xue, Jinyang Li, Zhengyuan Dong, Xinjian Zhao, Hao Xu, Chao Zhang, Reynold Cheng, M. Tamer \"Ozsu, Tianshu Yu
arXiv Machine Learning
Jul 15

Hierarchical Synthetic Tabular Data Generation: A Hybrid Top-Down and Bottom-Up Framework

arXiv:2605. 28198v2 Announce Type: replace Abstract: Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes.

By Junfeng Nie, Alvin Jin, Xiaohui Chen
arXiv AI
Aug 7

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

arXiv:2608. 06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard.

By Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem
Hugging Face Trending Papers
Aug 11

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur.

arXiv AI
Jun 26

Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.

By Mingxuan Jiang, Keyang Chen, Yongxin Wang, Yongsheng Zhao, Ziyue Dai, Yicun Liu, Zeping Li, Qiuyang Zhang, Hongyi Nie, Hongbin Zhu, Sen Liu, Guangnan Ye, Hongfeng Chai
arXiv AI
Jun 30

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

arXiv:2406. 08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate causal relationships that go beyond pairwise correlations largely unmeasured.

By Zineb Senane, Axel Karlsson, Lele Cao, Oleg Smirnov, Cheng Zhang, Sahar Asadi, Hedvig Kjellstr\"om, Gustav Eje Henter, Ruibo Tu