arXiv:2608. 14496v1 Announce Type: cross Abstract: Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets.
By Hao Yan, Lisa Pilgram, Dan Liu, Linglong Kong, Fida Dankar, Khaled El Emam
arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.
By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv:2403. 00965v2 Announce Type: replace-cross Abstract: Only a small fraction of patients with chronic kidney disease (CKD) progress to dialysis, creating severe class imbalance that limits the performance of machine learning models for early dialysis prediction.
By Hamed Khosravi, Milad Khanchi, Mobina Noori, Srinjoy Das, Abdullah Al-Mamun, Imtiaz Ahmed
The paper introduces a copula-based framework to relate Data‑Consistent Inversion (DCI) and its iterative variant (iDCI). By applying Sklar’s theorem, the authors factor the DCI update into marginal and dependence components, showing that any remaining discrepancy after iDCI convergence is fully captured by the copulas of the observed and predicted joint distributions. They prove that an exact copula transformation recovers the original DCI solution and provide convergence results for approximate transformations, supported by numerical examples illustrating adaptive refinement and progressive problem refinement.
By Troy Butler, Tianyi Jiang, Jo\~ao Silva, Harri Hakula, Timothy Wildey
arXiv:2606. 08903v1 Announce Type: new Abstract: Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity.
By Nicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm
MedFlow is a class‑aware multi‑scale flow matching framework designed to synthesize medical time‑series data. It uses a vector‑quantized multi‑scale tokenizer to capture both coarse and fine temporal patterns, and introduces Token Marginal Guidance to steer generation toward minority‑class characteristics. Experiments on four public datasets show MedFlow outperforms diffusion baselines, improving AUPRC by 5.8%, reducing Context‑FID by 88.6%, and achieving 3.8× higher sampling throughput.
By Yanhao Huang, Shibo Feng, Wanjin Feng, Peilin Zhao, Chunyan Miao