LeakageBench is a new benchmark comprising 500 document images with 11,954 GDPR‑aligned PII annotations, designed to evaluate document‑level redaction risk. It measures how well OCR pipelines, OCR‑dependent detectors, and OCR‑free vision‑language models can localize and remove sensitive information, using entity‑level F1, group‑wise leakage, and document‑level leakage metrics. The study shows that while advanced models improve localization, most pages still exhibit critical leakage, highlighting the need for higher‑recall, spatially grounded redaction methods.
By Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov
arXiv:2606. 18782v1 Announce Type: cross Abstract: Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII).
By Sean Brynj\'olfsson, Shashvat Jayakrishnan, Esha Sali, Diptanshu Purwar, Madhav Aggarwal
arXiv:2609.14352v1 Announce Type: new
Abstract: AI-generated image detection has attracted increasing attention, but existing evaluations mainly focus on natural images, leaving AI-generated document...
By Zhangjie Fu, Jiazhen Yan, Yuanwen Chen, Xinquan Yu, Yanzhe Li, Hui Jiang, Lei Gao, Chenfu Bao
arXiv:2606. 07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking external tools.
By Youting Wang, Yuan Tang, Yitian Qian, Chen Zhao
arXiv:2607. 16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information.
By Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding
T2LSC-Bench is a new benchmark for evaluating localized semantic control in text-to-image generation, consisting of 50 seed subjects and 1,200 prompt cases per model, producing 7,160 images across six models. The benchmark measures Text-at-Anchor Accuracy, Semantic Subject Preservation, Semantic Leakage Rate, and Conditional Semantic Leakage Rate using a dual‑branch protocol that combines OCR‑VLM verification with structured VLM semantic judgments. Results show that while accurate text rendering remains high, semantic leakage can increase dramatically under stress‑test conditions, and anti‑leakage prompting can reduce leakage without harming rendering accuracy.
By Yan Wang, Xinyi Hou, Weiguo Lin, Junjun Si, Siwei Ma