arXiv:2603. 10937v2 Announce Type: replace Abstract: The use of synthetic data has become increasingly popular as a privacy-preserving alternative to sharing real datasets, especially in sensitive domains such as healthcare, finance, and demography.
By Rajdeep Pathak, Amit Basak, Sayantee Jana
arXiv:2606. 00160v1 Announce Type: cross Abstract: Large language models (LLMs) suffer from degraded safety capabilities even when fine-tuned with benign datasets.
By Junbo Zhang, Qianli Zhou, Xinyang Deng, Wen Jiang, Jie Pan, Jinbiao Zhu
arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.
By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
The paper examines how supervised fine-tuning (SFT) of large language models can leak personally identifiable information (PII) when the fine-tuning data contains user-provided sensitive details. It introduces COVA, a coverage-aware decoding algorithm that improves targeted PII reconstruction from SFT models, especially when an adversary has limited contextual knowledge about a target. Experiments on medical and legal Q&A datasets show that even small proprietary SFT datasets can lead to significant privacy leakage via PII reconstruction.
By Sae Furukawa, Alina Oprea
arXiv:2602. 18518v2 Announce Type: replace Abstract: Content safety teams need metrics that reflect what users actually experience, not only what is reported.
By Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang, Faisal Farooq
arXiv:2607. 28862v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage.
By Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu
arXiv:2401.08038v2 Announce Type: replace-cross
Abstract: A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and com...
By Wenjun Qiu, David Lie, Lisa Austin
Calpric is a system that combines automatic text selection, segmentation, active learning, and crowdsourced annotation to create a large, balanced training set for privacy policy classification. By simplifying the labeling task, it enables untrained crowd workers to match the performance of trained annotators and reduces inter‑annotator disagreement, cutting labeling costs. The approach yields a dataset of 16,000 policy text segments across nine data categories and produces models that deliver accurate, fine‑grained labels at a cost of roughly $0.92–$1.71 per segment.
By Wenjun Qiu, David Lie, Lisa Austin
arXiv:2602.06268v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows. However, p...
By Junhyeok Lee, Han Jang, Kyu Sung Choi
arXiv:2605. 14746v2 Announce Type: replace Abstract: While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints.
By Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh, Yarin Gal, Yaniv Romano
The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.
By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
arXiv:2606. 09908v1 Announce Type: cross Abstract: Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evaluation.
By Ayana Hussain, Soumya Sharma, Golnoosh Farnadi, Nicholas Vincent, H\'eber Hwang Arcolezi, Ulrich A\"ivodji