Privacy-Preserving Deep Joint Source-Channel Coding with In-Loop Concept Erasure
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces a black-box membership inference attack framework tailored for fine-tuned text-to-speech models, addressing challenges in query generation and representation engineering. It evaluates five query types, finding recitation queries most effective, and uses multi-level speech embeddings with temporal alignment for fine-grained comparison. Experiments on CosyVoice2, F5-TTS, and XTTS-v2 trained on VCTK and British Dialect datasets show high privacy leakage, with speaker-level AUC above 0.80 and record-level AUC between 0.80 and 0.90.
The paper introduces VoxPrivacy, a benchmark for assessing interactional privacy in Speech Language Models (SLMs). It evaluates models on a 32‑hour bilingual dataset across three difficulty tiers, revealing that most open‑source SLMs perform near random on conditional privacy decisions and even strong closed‑source systems struggle with proactive privacy inference. The authors also validate these findings on a real‑speech subset and show that fine‑tuning on a 4,000‑hour training set can improve privacy‑preserving capabilities while maintaining robustness.
arXiv:2601. 17360v2 Announce Type: replace-cross Abstract: An adversary observing a model's released prediction can infer sensitive attributes of the queried input, or even reconstruct representatives of the model's training data.
The paper introduces a per-layer differential privacy (DP) clipping strategy for federated multilingual speech large language models (speech‑LLMs). It demonstrates that standard single‑pool per‑layer DP methods fail due to a cross‑component budget collapse caused by large norm differences between acoustic encoders and language decoders. The authors propose an α‑split two‑pool allocation that normalises encoder and decoder parameters separately, preserving the overall DP guarantee while restoring word error rate performance and providing tighter noise protection for the encoder.
arXiv:2609.36612v1 Announce Type: new Abstract: Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable conte...
The paper introduces DiffSem, a diffusion-based approach for task‑oriented semantic communications that splits the diffusion process between transmitter‑side self‑noising and receiver‑side reverse denoising. It addresses privacy concerns by reducing model‑inversion attacks while preserving task accuracy, as demonstrated on MNIST, CIFAR‑10, and CelebA datasets. The method achieves higher task performance without enlarging transmitted feature size or increasing semantic leakage.