arXiv:2508. 01725v5 Announce Type: replace Abstract: Recent advances in continuous conditional generative modeling, including Continuous conditional Generative Adversarial Network (CcGAN) and Continuous Conditional Diffusion Model (CCDM), estimate high-dimensional data distributions conditioned on scalar regression labels such as angles, ages, or temperatures.
By Xin Ding, Yun Chen, Yongwei Wang, Kao Zhang, Sen Zhang, Peibei Cao, Xiangxue Wang
EmbeddGAN introduces a new GAN framework that replaces the traditional discriminator with an embedding network trained to maximize statistical dependence between embeddings and real/fake labels using Gini distance correlation (gCor). The generator simultaneously minimizes this dependence, encouraging real and generated samples to become indistinguishable in the learned low‑dimensional embedding space. Experiments on MNIST, CIFAR‑10, and CelebA show competitive performance and notably more stable training dynamics compared to established baselines.
By MaTais Caldwell, Yixin Chen, Xin Dang, Charles Walter
arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.
By Mingxuan Jiang, Keyang Chen, Yongxin Wang, Yongsheng Zhao, Ziyue Dai, Yicun Liu, Zeping Li, Qiuyang Zhang, Hongyi Nie, Hongbin Zhu, Sen Liu, Guangnan Ye, Hongfeng Chai
arXiv:2509. 24935v3 Announce Type: replace-cross Abstract: Scalability has driven recent advances in generative modeling, yet its principles remain underexplored for adversarial learning.
By Sangeek Hyun, MinKyu Lee, Jae-Pil Heo
arXiv:2507. 14706v2 Announce Type: replace-cross Abstract: Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and the often subtle patterns that separate fraud from legitimate activity.
By Claudio Giusti, Luca Guarnera, Mirko Casu, Sebastiano Battiato
arXiv:2609.01410v1 Announce Type: cross
Abstract: Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly under...
By Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein
arXiv:2607. 16348v1 Announce Type: cross Abstract: Machine learning-based intrusion detection systems (IDSs) often suffer from class imbalance and vulnerability to adversarial attacks, leading to degraded detection performance and reduced robustness.
By Raihan Sultan Pasha Basuki, Aliyah Kurniasih
arXiv:2510.24046v2 Announce Type: replace-cross
Abstract: Existing tabular data generation methods primarily focus on matching statistical distributions between real and synthetic data, often overloo...
By Tu Anh Hoang Nguyen, Dang Nguyen, Tri-Nhan Vo, Thuc Duy Le, Trung Le, Sunil Gupta
arXiv:2405. 07332v2 Announce Type: cross Abstract: Numerous applications have resulted from the automation of agricultural disease segmentation using deep learning techniques.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Mohammad Shafiul Alam, Ahmed Al Wase, Md. Rabius Sani, Khan Md Hasib
arXiv:2608. 10096v1 Announce Type: cross Abstract: Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecified statistical models.
By Hyunjoo Kim, Sicheng Wu, Agastya Venkatraman, Guang Lin, Sehwan Kim
arXiv:2512.17730v2 Announce Type: replace
Abstract: Detectors of AI-generated images tend to inherit the biases of the data they are trained on: models fitted to GAN imagery learn to treat GAN-specif...
By Yichen Jiang, Mohammed Talha Alam, Sohail Ahmed Khan, Duc-Tien Dang-Nguyen, Fakhri Karray
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu