Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. However, Multi-modal Large Language Models (MLLMs), which process both text and images, introduce unique privacy challenges that remain underexplored.
The paper investigates whether inexpensive spectral metrics from the heavy‑tailed self‑regularisation framework can predict membership inference attack (MIA) vulnerability, offering a scalable alternative to costly shadow‑model attacks. Experiments on image and tabular classification tasks show that stable rank correlates positively with overall MIA success, while Log alpha‑Norm correlates negatively with MIA risk in low false‑positive regimes, outperforming conventional generalisation gap measures. These findings suggest that neural network spectra contain privacy leakage signals not captured by traditional overfitting metrics, pointing to spectral analysis as a promising direction for privacy auditing.
By Richard J. Preen, Jim Smith
arXiv:2606. 09125v1 Announce Type: cross Abstract: Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information.
By Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen, Hua Wei
arXiv:2608.28691v1 Announce Type: cross
Abstract: Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surround...
By Zhimin Li, Pan Wang, Jingxian Chen, Yuantao Tang, Anthony Chen, Qian Lou, Jingtong Hu
arXiv:2607. 02897v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns.
By Xianren Zhang, Delvin Ce Zhang, Dongwon Lee, Suhang Wang
QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.
By Shuaiqi Wang, Zinan Lin, Giulia Fanti