arXiv AI

Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification

arXiv:2606. 18372v1 Announce Type: cross Abstract: Educational dialogue is a valuable but sensitive resource for research: the same transcripts that capture authentic learning often capture personally identifiable information (PII) entangled with curricular content, where "Riemann" may refer to a real student or to a mathematical concept.

arXiv Machine Learning
Sep 14

Limits of LLM Text Detectors in Education

The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.

By Lukas Gehring, Benjamin Paa{\ss}en
arXiv Machine Learning
Sep 24

EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues

The paper introduces EduBehaviors, a framework that uses large language models to identify observable behaviors in educational dialogues and then trains a classifier to predict pedagogical constructs. By measuring repeated behaviors, the approach provides interpretable and scalable annotations, achieving macro‑F1 scores of 0.673 and Cohen’s kappa of 0.688 on the TalkMoves dataset. The authors also release the EduBehaviors Toolkit, enabling researchers to apply the framework to their own data.

By Julian Bernado, Ana Trindade Ribeiro, Xander Beberman, Susanna Loeb
arXiv AI
Jun 18

RedactionBench

arXiv:2606. 18782v1 Announce Type: cross Abstract: Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII).

By Sean Brynj\'olfsson, Shashvat Jayakrishnan, Esha Sali, Diptanshu Purwar, Madhav Aggarwal
arXiv AI
Jun 10

IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts

arXiv:2606. 09908v1 Announce Type: cross Abstract: Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evaluation.

By Ayana Hussain, Soumya Sharma, Golnoosh Farnadi, Nicholas Vincent, H\'eber Hwang Arcolezi, Ulrich A\"ivodji
arXiv Machine Learning
Aug 27

Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models

The paper examines how supervised fine-tuning (SFT) of large language models can leak personally identifiable information (PII) when the fine-tuning data contains user-provided sensitive details. It introduces COVA, a coverage-aware decoding algorithm that improves targeted PII reconstruction from SFT models, especially when an adversary has limited contextual knowledge about a target. Experiments on medical and legal Q&A datasets show that even small proprietary SFT datasets can lead to significant privacy leakage via PII reconstruction.

By Sae Furukawa, Alina Oprea
arXiv AI
Aug 20

Redakto - The Incognito Tab for LLMs

Redakto is a new tool designed to anonymize text before it is processed by large language models (LLMs). It offers state‑of‑the‑art redaction of personally identifiable information (PII) and pseudonymization, accessible via a web interface, REST APIs, and model context protocol hooks. The authors evaluate its performance on legal and medical datasets, showing that anonymized texts retain utility comparable to the originals, enabling LLM tasks without significant loss of effectiveness.

By Saurav Kumar Saha, Tom R\"ohr, Felix Bie{\ss}mann
arXiv AI
Sep 25

ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction

ASIRF (Agentic Sensitive Information Redaction Framework) is a system that retrieves domain‑specific definitions of sensitive information from a flexible knowledge base at inference time, eliminating the need for retraining when adapting to new domains. It offers two architectures—a three‑call multi‑agent pipeline and a single‑agent variant—and has been evaluated on ten small open‑weight models across eight datasets, including out‑of‑distribution fictional domains. In 68 of 80 model‑domain combinations (85 %), ASIRF’s recall surpasses that of the OpenAI Privacy Filter, with most shortfalls limited to the filter’s training‑distribution domains.

By Sudha Priyadarshini, Mohamed Chahine Ghanem