Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech
Read the original on arXiv AI →The paper introduces GUARD, a lightweight speaker identity unlearning framework designed to prevent re-identification in zero-shot text-to-speech systems. GUARD employs a learned speaker gate and speaker-agnostic activation steering on a frozen TTS backbone, optimizing steering vectors through group-relative reward optimization to reduce similarity to forgotten speakers while maintaining intelligibility and naturalness. Experiments on CosyVoice2 show that GUARD significantly lowers forget-speaker similarity and re-identification accuracy while preserving the ability to reproduce retained speakers.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.