arXiv:2608. 14496v1 Announce Type: cross Abstract: Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets.
By Hao Yan, Lisa Pilgram, Dan Liu, Linglong Kong, Fida Dankar, Khaled El Emam
arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.
By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv:2403. 00965v2 Announce Type: replace-cross Abstract: Only a small fraction of patients with chronic kidney disease (CKD) progress to dialysis, creating severe class imbalance that limits the performance of machine learning models for early dialysis prediction.
By Hamed Khosravi, Milad Khanchi, Mobina Noori, Srinjoy Das, Abdullah Al-Mamun, Imtiaz Ahmed
The paper introduces a copula-based framework to relate Data‑Consistent Inversion (DCI) and its iterative variant (iDCI). By applying Sklar’s theorem, the authors factor the DCI update into marginal and dependence components, showing that any remaining discrepancy after iDCI convergence is fully captured by the copulas of the observed and predicted joint distributions. They prove that an exact copula transformation recovers the original DCI solution and provide convergence results for approximate transformations, supported by numerical examples illustrating adaptive refinement and progressive problem refinement.
By Troy Butler, Tianyi Jiang, Jo\~ao Silva, Harri Hakula, Timothy Wildey
arXiv:2606. 08903v1 Announce Type: new Abstract: Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity.
By Nicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm
MedFlow is a class‑aware multi‑scale flow matching framework designed to synthesize medical time‑series data. It uses a vector‑quantized multi‑scale tokenizer to capture both coarse and fine temporal patterns, and introduces Token Marginal Guidance to steer generation toward minority‑class characteristics. Experiments on four public datasets show MedFlow outperforms diffusion baselines, improving AUPRC by 5.8%, reducing Context‑FID by 88.6%, and achieving 3.8× higher sampling throughput.
By Yanhao Huang, Shibo Feng, Wanjin Feng, Peilin Zhao, Chunyan Miao
Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity. We introduce a multi-dimensional evaluation framework grounded in epidemiology, assessing descriptive fidelity, clinical utility, and structural validity, corresponding to descriptive, predictive, and causal questions.
arXiv:2608.21673v1 Announce Type: cross
Abstract: Longitudinal electronic health records (EHRs) document patients' sequences of clinical visits over time, preserving the temporal evolution of disease...
By Ximiao Li, Lin Jiang, Rongchao Xu, Dahai Yu, Zhe He, Guang Wang
The paper presents Copula Adapted Directed Acyclic Graph (CopDAG), a framework that combines copula models with an ensemble of causal structure discovery methods based on Directed Acyclic Graphs to represent biomedical data. By capturing non‑Gaussian, non‑linear dependencies and stable causal relationships, CopDAG enables clustering of unlabeled biomedical data using K‑means. Across 16 biomedical datasets, CopDAG achieves the highest normalized clustering accuracy and adjusted Rand index among 12 evaluated methods, and it can predict class labels and provide explainable causal visualizations without relying on data annotations.
By Heranga K. Rathnasekara, Norou Diawara, Manar D. Samad
arXiv:2606. 06990v1 Announce Type: new Abstract: The generation of high-fidelity synthetic Electronic Health Records (EHR) is crucial for advancing medical research while preserving patient privacy.
By Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie, Jimeng Sun
arXiv:2605.23632v2 Announce Type: replace
Abstract: We introduce Gaussian Mixture Copula Processes (GMCP), a conditional copula process for irregularly sampled multivariate time series (IMTS) that is...
By Christian Kl\"otergens, Tom Hanika, Lars Schmidt-Thieme, Vijaya Krishna Yalavarthi
arXiv:2607. 25020v1 Announce Type: new Abstract: Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas.
By Nicholas Andrea Pearson, Francesca Zanello, Davide Russo, Luca Bortolussi, Francesca Cairoli