arXiv:2608. 14496v1 Announce Type: cross Abstract: Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets.
By Hao Yan, Lisa Pilgram, Dan Liu, Linglong Kong, Fida Dankar, Khaled El Emam
arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.
By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv:2606. 19643v1 Announce Type: cross Abstract: Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.
By Julie Fendler, Francesca L. Crowe, Tom Marshall, Sylvia Richardson, Paul D. W. Kirk
arXiv:2603. 23016v2 Announce Type: replace-cross Abstract: Tabular data is more challenging to generate than text and images, due to its heterogeneous features and much lower sample sizes.
By Davide Scassola, Dylan Ponsford, Adri\'an Javaloy, Sebastiano Saccani, Luca Bortolussi, Henry Gouk, Antonio Vergari
arXiv:2609.39124v1 Announce Type: new
Abstract: Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many spe...
By Mohamed Amine Ketata, Maximilian Schambach, Stephan G\"unnemann
arXiv:2609.00773v1 Announce Type: cross
Abstract: High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t...
By Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan
GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv:2609.01433v1 Announce Type: new
Abstract: Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semant...
By Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng, Zhengchun Zhou
arXiv:2609.22654v1 Announce Type: cross
Abstract: Federated learning (FL) has emerged as a leading privacy-preserving framework for collaborative machine learning across decentralized environments. W...
By Brigham Halverson, Sharmistha Guha, Jessica Bernard, Rajarshi Guhaniyogi
The paper introduces the General Demographic Pre-trained (GDP) model, a lightweight foundation model that learns representations from the two most common clinical attributes—age and sex. By optimizing encoding and visit‑reordering strategies, GDP embeddings are shown to improve predictive performance when concatenated with raw features across various disease and geographic cohorts. The model outperforms several state‑of‑the‑art tabular foundation models and tree‑based algorithms, demonstrating that enriched demographic embeddings can enhance classification tasks while remaining fully compatible with standard classifiers.
By Li-Chin Chen, Ji-Tian Sheu, Yuh-Jue Chuang
arXiv:2609.16069v1 Announce Type: cross
Abstract: Synthetic tabular data can match real data distributions while still violating the semantic constraints that govern valid tabular rows. This reveals...
By Yili Wang, Ruxue Shi, Mengnan Du, Hangting Ye, Yi Chang, Xin Wang
arXiv:2607. 01275v1 Announce Type: cross Abstract: Variational Autoencoders (VAEs) commonly assume a standard isotropic Gaussian prior over the latent space, an assumption that often fails to capture the true distribution of latent representations for complex datasets.
By Qijun Chen, Shaofan Li