arXiv:2608. 03174v1 Announce Type: cross Abstract: Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identifying model-generated outputs.
By Miryam Mi-Ying Huang, Chung-Wei Lee, Max Raffel, Er-Cheng Tang
arXiv:2606. 11698v1 Announce Type: cross Abstract: Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures.
By Jian-Ping Mei, Weibin Zhang, Ao Yao, Tiantian Zhu, Jie Xiao
arXiv:2607. 10554v1 Announce Type: cross Abstract: With the development of generative AI, watermarking techniques have been widely used to detect the authenticity of AI-generated data and protect the rights of users and creators.
By Dongyu Cui, Xuan Bi
arXiv:2607. 00325v1 Announce Type: new Abstract: A growing body of literature suggests that training data membership inference problems are fundamentally hard tasks in modern language modeling settings.
By John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Tom Goldstein
arXiv:2608. 19727v1 Announce Type: cross Abstract: Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks.
By Dongbin Kim, Geonwoo Shin, Yujin Choi, Soyeon Park, Jaewook Lee
arXiv:2609.39024v1 Announce Type: new
Abstract: Text-to-image (T2I) generation is gaining increasing popularity with the general public, motivating the development of reliable mechanisms for copyrigh...
By Dixi Yao, Kaiwen Chen, Tahseen Rabbani, Tian Li
arXiv:2608.20580v1 Announce Type: cross
Abstract: Federated learning (FL) is vulnerable to multi-level attacks. However, existing methods address them separately, leaving FL exposed to data leakage,...
By Xinyun Liu, Zhi Lu, Yu Chen, Ronghua Xu
arXiv:2608. 10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored.
By Jie Cao, Qi Li, Zelin Zhang, Xiaodong Wu, Lingshuang Liu, Xiangman Li, Jianbing Ni
arXiv:2409. 06130v2 Announce Type: replace-cross Abstract: Modern machine learning models require substantial computational resources and data to train, making them valuable intellectual property.
By Aoting Hu, Yanzhi Chen, Renjie Xie, Xinwei Zhang, Wei Xu
arXiv:2603. 10937v2 Announce Type: replace Abstract: The use of synthetic data has become increasingly popular as a privacy-preserving alternative to sharing real datasets, especially in sensitive domains such as healthcare, finance, and demography.
By Rajdeep Pathak, Amit Basak, Sayantee Jana
arXiv:2607. 13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT).
By Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu
arXiv:2608.29674v1 Announce Type: new
Abstract: Sharing tabular data in high-stakes domains is constrained by privacy regulations. Synthetic data offer a promising alternative, but deep generative mo...
By Jinmeng Li, Quan Zhang, Hangting Ye, He Zhao, Firas Laakom, Dandan Guo, J\"urgen Schmidhuber