arXiv AI By Duong Bach, Hai Nguyen Hong, Cuong Do

Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models

Read the original on arXiv AI →

arXiv:2608. 05243v1 Announce Type: cross Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 20

GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

GEAR is a two‑stage framework that distills tabular foundation models into lightweight MLP or tree‑based predictors for efficient CPU deployment. In the first stage, synthetic covariates are used as teacher‑query locations to train the student on soft TFM targets, expanding coverage beyond observed rows. The second stage re‑anchors the student to the target distribution using real labels and out‑of‑fold teacher predictions, preventing self‑labeling leakage and improving performance. Experiments on TALENT and TabArena show that GEAR‑distilled MLPs outperform supervised MLPs by up to 2.00 AUC points on binary tasks and 1.35 on multiclass tasks, and also outperform CatBoost, while dramatically reducing inference time and memory usage.

By Qi Qin, Jiajie Zhu, Dali Chen, Yuzhao Zhang, Jia-Xing Han, Yu Su, Peng Zhang, Ying Yan, Yifan Sun
arXiv AI
Jun 24

MGI: Member vs Generated Inference

arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.

By Bihe Zhao, Michel Meintz, Juangui Xu, Franziska Boenisch, Adam Dziedzic
arXiv Machine Learning
Sep 10

Selective Posterior Margin Regularization for Forward-Corrected Classification

The paper introduces Selective Posterior Margin Regularization (SPMR), a technique that enhances Forward correction for learning with class‑conditional label noise. SPMR preserves the Forward objective while converting disagreements between the corrected likelihood’s reverse posterior and the observed annotation into a graded update on the clean classifier. Experiments on five known‑transition benchmarks show that SPMR improves performance by 2.5–7.0 percentage points over full‑length Forward and remains 0.7–2.5 percentage points better when combined with Mixup and early stopping, with gains attributed to posterior‑space coefficients, transition‑adjusted targets, and pairwise actions.

By Zexing Zhang, Jichao Li, Tianyang Lei, XiongYi Lu, Yang Kewei