arXiv:2609.22112v1 Announce Type: new
Abstract: Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that...
By Muhammed Nazmul Arefin, Omar Jamal Hammad
arXiv:2608.22161v1 Announce Type: new
Abstract: Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existi...
By Qian Ma, Anna Squicciarini, Sarah Rajtmajer
arXiv:2601. 04641v2 Announce Type: replace-cross Abstract: The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation.
By Lionel Z. Wang, Yusheng Zhao, Jiabin Luo, Xinfeng Li, Lixu Wang, Yinan Peng, Haoyang Li, XiaoFeng Wang, Wei Dong
The paper "Evaluating Style-Personalized Text Generation: Challenges and Directions" examines the difficulties of assessing text that is tailored to individual users’ styles. It critiques common metrics such as BLEU, embeddings, and LLM-as-judges, and introduces a style discrimination benchmark covering domain discrimination, authorship attribution, and LLM-generated personalized versus non-personalized discrimination across eight writing tasks. The study finds that ensembles of diverse evaluation metrics outperform single-evaluator approaches and offers guidance for reliable assessment of style-personalized generation.
By Anubhav Jangra, Bahareh Sarrafzadeh, Silviu Cucerzan, Adrian de Wynter, Sujay Kumar Jauhar
arXiv:2505. 14608v3 Announce Type: replace-cross Abstract: Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be manipulated to evade detection has led to suggestions that the problem is inherently intractable.
By Rafael Rivera Soto, Barry Chen, Nicholas Andrews
arXiv:2606. 10099v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns about misuse such as plagiarism, misinformation, and automated influence operations, motivating the need for robust detectors.
By Rafael Rivera Soto, Barry Chen, Nicholas Andrews
arXiv:2510.13302v4 Announce Type: replace-cross
Abstract: Computational stylometry studies writing style through quantitative textual patterns, enabling applications such as authorship attribution, i...
By Pablo Miralles-Gonz\'alez, Javier Huertas-Tato, Alejandro Mart\'in, David Camacho
SpeechLLMs used in professional settings often undergo domain customisation through prompts or fine‑tuning, which can inadvertently cause the model to transcribe phonetically similar words from its context or training data, leaking private information. The authors systematically investigate this overlooked privacy risk, creating benchmarks to measure leakage rates for both prompting and fine‑tuning, and find that both mechanisms cause measurable leakage that compounds when combined. They evaluate a prompt‑level mitigation strategy and analyse the accuracy‑leakage trade‑off, concluding that fine‑tuning without context prompts offers the best balance between performance and privacy.
By Maike Z\"ufle, Jan Niehues
HyperStyler is a new architecture for low-resource authorship style transfer that separates style selection from style realization. It uses a style navigator to predict style coordinates from source context and target-author references, and a style hypernetwork to apply these coordinates through dynamic parameter modulation. Experiments on Reddit, Blog, and News datasets show that HyperStyler outperforms previous methods, including LLM-based approaches, while adding only 2.4% more parameters than T5-large and running 1.8× faster at inference.
By Jongkyung Shin, Minguk Jeon, Chanwoo Park, Chiehyeon Lim
The paper introduces a two‑stage speech anonymization framework that preserves both linguistic content and acoustic identity. It replaces personally identifiable information using a generative editing model and applies a flow‑matching anonymization technique (F3‑VA) to create diverse, distinct anonymized speakers. The authors evaluate privacy with speaker verification metrics and utility by training ASR, TTS, and SER models from scratch, showing stronger privacy protection with minimal utility loss compared to existing baselines.
By Yunchong Xiao, Yuxiang Zhao, Ziyang Ma, Shuai Wang, Kai Yu, Jiachun Liao, Xie Chen
The paper introduces the privacy‑HSD trade‑off, highlighting that automatic hate speech detection systems can inadvertently compromise user privacy by encoding authorship. It demonstrates that such systems may achieve high performance at the expense of privacy, and proposes a new domain‑specific technique, AgnoSpeech, alongside other text privatization methods to balance these competing goals. The authors benchmark these methods, showing that while challenging, it is feasible to protect privacy without sacrificing hate‑speech detection effectiveness.
arXiv:2603.15220v2 Announce Type: replace
Abstract: Strict anonymity of model responses is a key for the reliability of voting-based leaderboards, such as LM Arena. While prior studies have attempted...
By Minsung Cho, Jaehyung Kim