arXiv Machine Learning By Laure Berti-Equille

Model-Aware Data Cleaning for Tabular Foundation Models

Read the original on arXiv Machine Learning →

The paper introduces L2C‑TFM, a reinforcement‑learning framework for cleaning tabular data before feeding it to Tabular Foundation Models (TFMs). It proposes a model‑aware reward that regularizes the Wasserstein distance between cleaned and dirty data, aiming to preserve distributional stability. Experiments on ten OpenML datasets show that while some reward designs fail, the model‑aware reward performs comparably to a random‑forest baseline and improves minority‑class macro‑F1 under class imbalance, and a policy trained on one dataset can transfer to others.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 18

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

arXiv:2606. 19162v1 Announce Type: new Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself.

By Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana Romero-Soriano, Michal Drozdzal
arXiv Machine Learning
4d ago

TabFM: A Zero-Shot Foundation Model for Tabular Data

arXiv:2609.37959v1 Announce Type: new Abstract: Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We p...

By Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das