arXiv Computation and Language By Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov

LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images

Read the original on arXiv Computation and Language →

LeakageBench is a new benchmark comprising 500 document images with 11,954 GDPR‑aligned PII annotations, designed to evaluate document‑level redaction risk. It measures how well OCR pipelines, OCR‑dependent detectors, and OCR‑free vision‑language models can localize and remove sensitive information, using entity‑level F1, group‑wise leakage, and document‑level leakage metrics. The study shows that while advanced models improve localization, most pages still exhibit critical leakage, highlighting the need for higher‑recall, spatially grounded redaction methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Sep 2

LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images

LeakageBench is a new benchmark consisting of 500 document images with 11,954 GDPR‑aligned PII annotations, designed to evaluate document‑level redaction risk. Unlike existing text‑centric PII benchmarks, it measures whether a page remains unsafe if any identifier is missed, using entity‑level F1, group‑wise leakage, and document‑level leakage metrics. Experiments show that even advanced OCR pipelines and vision‑language models improve localization but still leave a high proportion of pages unsafe for release.

arXiv AI
Jun 18

RedactionBench

arXiv:2606. 18782v1 Announce Type: cross Abstract: Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII).

By Sean Brynj\'olfsson, Shashvat Jayakrishnan, Esha Sali, Diptanshu Purwar, Madhav Aggarwal
arXiv AI
Jul 21

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.

By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
arXiv Machine Learning
Aug 19

Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification

The paper introduces Strategic 16K, a 16,000‑document corpus of diplomatic cables from WikiLeaks’ Public Library of US Diplomacy, designed to eliminate label leakage. It benchmarks six models—both classical machine learning and transformer-based—on this leakage‑controlled dataset, finding BERT and ELECTRA top performers while TF‑IDF with Logistic Regression offers strong accuracy at lower cost. This work provides the first fully reproducible sensitivity‑classification benchmark built under explicit leakage‑control conditions.

By Aleesha Zainab, Muhammad Ahmed Khalid, Faheem Ullah Khan, Asifullah Khan