arXiv Machine Learning

Loss Landscape Poisoning: Targeted Extraction of Unseen Training Data from LLMs

arXiv:2606. 17110v1 Announce Type: cross Abstract: Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user conversations containing secrets.

arXiv Machine Learning
Jul 16

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

arXiv:2607. 13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT).

By Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu