arXiv Machine Learning By Ayeen Poostforoushan, Liane Vogel, Carsten Binnig

TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings

Read the original on arXiv Machine Learning →

arXiv:2607. 24130v1 Announce Type: cross Abstract: Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

The paper introduces TEmBed, a unified benchmark for evaluating tabular embeddings across four representation levels—cell, row, column, and table—using a diverse set of models. It demonstrates that the best model depends on the specific task and representation level, providing practical guidance for selecting embeddings in real-world applications. The study aims to facilitate the development of more general-purpose tabular representation models.

By Liane Vogel, Kavitha Srinivas, Niharika D'Souza, Sola Shirai, Oktie Hassanzadeh, Horst Samulowitz
arXiv AI
Jun 9

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

arXiv:2606. 09323v1 Announce Type: new Abstract: Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare directly even when they operate on similar tabular signals.

By Wei Pang, Xiangru Jian, Hehan Li, Zhixuan Yu, Alex Xue, Jinyang Li, Zhengyuan Dong, Xinjian Zhao, Hao Xu, Chao Zhang, Reynold Cheng, M. Tamer \"Ozsu, Tianshu Yu
arXiv AI
Jun 30

Beyond IID: How General Are Tabular Foundation Models, Really?

arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.

By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
arXiv Machine Learning
Aug 4

LakeMLB: Data Lake Machine Learning Benchmark

arXiv:2602. 10441v2 Announce Type: replace Abstract: Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data.

By Feiyu Pan, Tianbin Zhang, Aoqian Zhang, Yu Sun, Zheng Wang, Lixing Chen, Li Pan, Jianhua Li