arXiv:2609.20999v1 Announce Type: cross
Abstract: Latent variable generative models are commonly fit using simple priors over latent variables, but draws from these priors often fail to produce reali...
By Shweta Dutta, Gemma E. Moran
GEAR is a two‑stage framework that distills tabular foundation models into lightweight MLP or tree‑based predictors for efficient CPU deployment. In the first stage, synthetic covariates are used as teacher‑query locations to train the student on soft TFM targets, expanding coverage beyond observed rows. The second stage re‑anchors the student to the target distribution using real labels and out‑of‑fold teacher predictions, preventing self‑labeling leakage and improving performance. Experiments on TALENT and TabArena show that GEAR‑distilled MLPs outperform supervised MLPs by up to 2.00 AUC points on binary tasks and 1.35 on multiclass tasks, and also outperform CatBoost, while dramatically reducing inference time and memory usage.
By Qi Qin, Jiajie Zhu, Dali Chen, Yuzhao Zhang, Jia-Xing Han, Yu Su, Peng Zhang, Ying Yan, Yifan Sun
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.
By Bihe Zhao, Michel Meintz, Juangui Xu, Franziska Boenisch, Adam Dziedzic
arXiv:2607. 21636v1 Announce Type: new Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and clinical risk.
By Jie Zhang
The paper introduces Selective Posterior Margin Regularization (SPMR), a technique that enhances Forward correction for learning with class‑conditional label noise. SPMR preserves the Forward objective while converting disagreements between the corrected likelihood’s reverse posterior and the observed annotation into a graded update on the clean classifier. Experiments on five known‑transition benchmarks show that SPMR improves performance by 2.5–7.0 percentage points over full‑length Forward and remains 0.7–2.5 percentage points better when combined with Mixup and early stopping, with gains attributed to posterior‑space coefficients, transition‑adjusted targets, and pairwise actions.
By Zexing Zhang, Jichao Li, Tianyang Lei, XiongYi Lu, Yang Kewei
arXiv:2607. 24943v1 Announce Type: cross Abstract: In many classification problems, reliable instance-level labels are unavailable.
By Rapha\"el Bonnet-Guerrini, Johann Ioannou-Nikolaides, Troels Petersen, Vincenzo Piuri
BLADE is a variational model for knowledge graph completion that separates latent truth from graph recording and uses distilled offline language‑model judgments as a frozen teacher regularizer. The model provides calibrated probabilities and epistemic uncertainty through posterior samples, while the teacher is only an optional triage factor during inference. Across five benchmarks, BLADE matches ranking performance and significantly reduces expected calibration error, improving ECE, Brier score, and NLL over several baselines, and shows strong performance under controlled missingness and leakage stress tests.
By Ibne Farabi Shihab, Rabeya Bosri Tamanna, Abdo El Karaky, Sanjeda Akter, Anuj Sharma
arXiv:2405. 07780v3 Announce Type: replace-cross Abstract: This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced.
By Zhiyong Yang, Qianqian Xu, Sicong Li, Zitai Wang, Xiaochun Cao, Qingming Huang
arXiv:2608.28931v1 Announce Type: cross
Abstract: Matching users to interest categories at scale is central to personalized shopping, but the task is challenging in large e-commerce platforms, where...
By Abhinav Mahajan, Arindam Sarkar, Prakash Mandayam Comar
arXiv:2606. 21806v2 Announce Type: replace Abstract: Deep generative models reproduce the observational distribution of their training data, inheriting any spurious associations it contains.
By Jingyuan Chen, Kangrui Ruan, Junzhe Zhang
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu