arXiv:2606. 07996v1 Announce Type: cross Abstract: Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and raises ethical, legal, and fairness concerns.
By Kaixin Lan, Mu You, Tao Fang, Binkai Ou, Lidia S. Chao, Derek F. Wong
arXiv:2604. 01904v3 Announce Type: replace-cross Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.
By Muxing Li, Zesheng Ye, Sharon Li, Feng Liu
arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.
By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
Redakto is a new tool designed to anonymize text before it is processed by large language models (LLMs). It offers state‑of‑the‑art redaction of personally identifiable information (PII) and pseudonymization, accessible via a web interface, REST APIs, and model context protocol hooks. The authors evaluate its performance on legal and medical datasets, showing that anonymized texts retain utility comparable to the originals, enabling LLM tasks without significant loss of effectiveness.
By Saurav Kumar Saha, Tom R\"ohr, Felix Bie{\ss}mann
arXiv:2607. 22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment.
By Ryan Thornton, Mir Mehedi Ahsan Pritom, Maanak Gupta
The paper surveys how large language models (LLMs) are being applied in legal tasks such as judgement prediction, document analysis, and drafting. It reviews the benefits of automation while highlighting legal challenges like privacy, bias, and explainability. The authors also discuss data resources for legal domain specialization and outline future research directions.
By Zhongxiang Sun
The paper demonstrates that fine‑tuning large language models on a single author’s works can trigger the models to reproduce large verbatim excerpts from copyrighted books, even when prompted only with semantic descriptions. Experiments on GPT‑4o, Gemini‑2.5‑Pro, and DeepSeek‑V3.1 show up to 85‑90% recall of held‑out books, with spans exceeding 460 words, and this effect generalizes across authors and model providers. The findings suggest that fine‑tuning reactivates latent memorization from pre‑training, revealing a widespread vulnerability in industry models.
By Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg, Tuhin Chakrabarty
arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.
By Lorenzo Rossi, Bart{\l}omiej Marek, Franziska Boenisch, Adam Dziedzic
Deep Contrastive Unlearning for Language Models (DeepCUT) is a framework that removes information from fine‑tuned language models by directly optimizing their latent space. It addresses the challenge of machine unlearning in black‑box models, which has been largely overlooked by previous work that only mitigated output effects. Experiments on real‑world datasets show that DeepCUT consistently outperforms baseline methods in both effectiveness and efficiency.
By Estrid He, Tabinda Sarwar, Ibrahim Khalil, Xun Yi, Ke Wang
arXiv:2502.10673v2 Announce Type: replace-cross
Abstract: Retrieval-Augmented Generation (RAG) has become an effective method for enhancing large language models (LLMs) with up-to-date knowledge. How...
By Yepeng Liu, Xuandong Zhao, Dawn Song, Yuheng Bu
arXiv:2602. 18733v2 Announce Type: replace Abstract: Training data leakage from Large Language Models (LLMs) raises serious concerns related to privacy, security, and copyright compliance.
By Trishita Tiwari, Ari Trachtenberg, G. Edward Suh
arXiv:2606. 16244v1 Announce Type: cross Abstract: Large language models routinely generate code with exploitable security flaws.
By Xiaoyun Xu, Lichao Wu, Jona te Lintelo, Siyu Zhang, Stjepan Picek