arXiv AI

Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path

arXiv:2606. 07271v1 Announce Type: cross Abstract: Understanding what generative models retain from training data remains challenging, with implications for copyright and privacy.

arXiv AI
Jun 24

MGI: Member vs Generated Inference

arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.

By Bihe Zhao, Michel Meintz, Juangui Xu, Franziska Boenisch, Adam Dziedzic
arXiv Machine Learning
Jul 16

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

arXiv:2607. 13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT).

By Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu
arXiv AI
Jun 2

Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization

arXiv:2510. 10982v2 Announce Type: replace-cross Abstract: Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications.

By Zihan Wang, Zhiyong Ma, Zhongkui Ma, Shuofeng Liu, Akide Liu, Derui Wang, Minhui Xue, Guangdong Bai
arXiv AI
4d ago

Calibrating One-Round Membership Inference with Neighbors

The paper addresses the challenge of calibrating membership inference attacks in a one‑round setting where only a single trained model is available. It proposes using neighboring data points of the target to approximate the calibration that reference models normally provide, and demonstrates that querying these neighbors—especially against early training checkpoints—enhances the membership signal. Experiments on three image classification datasets and training setups show that this neighbor‑based approach yields strong attack performance without extra training cost.

By Francesco Rita, Jie Zhang, Florian Tram\`er
arXiv Machine Learning
Jun 5

Zero-Flow Encoders

arXiv:2602. 00797v2 Announce Type: replace-cross Abstract: Flow-based methods have achieved significant success in various generative modeling tasks, capturing nuanced details within complex data distributions.

By Yakun Wang, Leyang Wang, Song Liu, Taiji Suzuki
arXiv Computer Vision
Aug 27

DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.

By Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu