arXiv:2603. 02150v2 Announce Type: replace-cross Abstract: The extraction of critical information from crime-related documents is a crucial task for law enforcement agencies.
By Miguel Lopez-Duran, Julian Fierrez, Aythami Morales, Daniel DeAlcala, Gonzalo Mancera, Javier Irigoyen, Ruben Tolosana, Oscar Delgado, Francisco Jurado, Alvaro Ortigosa
arXiv:2606. 19710v1 Announce Type: cross Abstract: Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargon-heavy legal documents.
By Elijah Feldman, Dipak Meher, Carlotta Domeniconi
arXiv:2609.22529v1 Announce Type: new
Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its t...
By Genis Skura, Roland Bouffanais, Didier Wernli
The paper introduces a scalable, multi-step framework designed to improve the quality of Named Entity Recognition (NER) annotations, particularly in low-resource languages. It employs a frequency-based iterative approach that combines self‑training with a dual‑threshold mechanism to increase inference confidence. Experiments on various NER datasets show notable performance gains over the original data, and the study also investigates the use of generative Large Language Models for NER tasks.
By Toqeer Ehsan, Thamar Solorio
arXiv:2606. 19852v1 Announce Type: cross Abstract: Information extraction from pathology reports is essential for cancer staging, tumor registry population.
By Aman Pathak, Cheng Peng, Mengxian Lyu, Ziyi Chen, Reema Solan, Sankalp Talankar, Yasir Khan, Hiren Mehta, Aokun Chen, Yi Guo, Yonghui Wu
The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially in inferential categories like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.