LeakageBench is a new benchmark consisting of 500 document images with 11,954 GDPR‑aligned PII annotations, designed to evaluate document‑level redaction risk. Unlike existing text‑centric PII benchmarks, it measures whether a page remains unsafe if any identifier is missed, using entity‑level F1, group‑wise leakage, and document‑level leakage metrics. Experiments show that even advanced OCR pipelines and vision‑language models improve localization but still leave a high proportion of pages unsafe for release.
arXiv:2609.14352v1 Announce Type: new
Abstract: AI-generated image detection has attracted increasing attention, but existing evaluations mainly focus on natural images, leaving AI-generated document...
By Zhangjie Fu, Jiazhen Yan, Yuanwen Chen, Xinquan Yu, Yanzhe Li, Hui Jiang, Lei Gao, Chenfu Bao
arXiv:2606. 18782v1 Announce Type: cross Abstract: Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII).
By Sean Brynj\'olfsson, Shashvat Jayakrishnan, Esha Sali, Diptanshu Purwar, Madhav Aggarwal
arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.
By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
arXiv:2606. 07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking external tools.
By Youting Wang, Yuan Tang, Yitian Qian, Chen Zhao
The paper introduces Strategic 16K, a 16,000‑document corpus of diplomatic cables from WikiLeaks’ Public Library of US Diplomacy, designed to eliminate label leakage. It benchmarks six models—both classical machine learning and transformer-based—on this leakage‑controlled dataset, finding BERT and ELECTRA top performers while TF‑IDF with Logistic Regression offers strong accuracy at lower cost. This work provides the first fully reproducible sensitivity‑classification benchmark built under explicit leakage‑control conditions.
By Aleesha Zainab, Muhammad Ahmed Khalid, Faheem Ullah Khan, Asifullah Khan