arXiv:2609.20999v1 Announce Type: cross
Abstract: Latent variable generative models are commonly fit using simple priors over latent variables, but draws from these priors often fail to produce reali...
By Shweta Dutta, Gemma E. Moran
GEAR is a two‑stage framework that distills tabular foundation models into lightweight MLP or tree‑based predictors for efficient CPU deployment. In the first stage, synthetic covariates are used as teacher‑query locations to train the student on soft TFM targets, expanding coverage beyond observed rows. The second stage re‑anchors the student to the target distribution using real labels and out‑of‑fold teacher predictions, preventing self‑labeling leakage and improving performance. Experiments on TALENT and TabArena show that GEAR‑distilled MLPs outperform supervised MLPs by up to 2.00 AUC points on binary tasks and 1.35 on multiclass tasks, and also outperform CatBoost, while dramatically reducing inference time and memory usage.
By Qi Qin, Jiajie Zhu, Dali Chen, Yuzhao Zhang, Jia-Xing Han, Yu Su, Peng Zhang, Ying Yan, Yifan Sun
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.
By Bihe Zhao, Michel Meintz, Juangui Xu, Franziska Boenisch, Adam Dziedzic
arXiv:2607. 21636v1 Announce Type: new Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and clinical risk.
By Jie Zhang
The paper introduces Selective Posterior Margin Regularization (SPMR), a technique that enhances Forward correction for learning with class‑conditional label noise. SPMR preserves the Forward objective while converting disagreements between the corrected likelihood’s reverse posterior and the observed annotation into a graded update on the clean classifier. Experiments on five known‑transition benchmarks show that SPMR improves performance by 2.5–7.0 percentage points over full‑length Forward and remains 0.7–2.5 percentage points better when combined with Mixup and early stopping, with gains attributed to posterior‑space coefficients, transition‑adjusted targets, and pairwise actions.
By Zexing Zhang, Jichao Li, Tianyang Lei, XiongYi Lu, Yang Kewei