arXiv AI

CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

arXiv:2602. 17663v3 Announce Type: replace Abstract: HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts.

arXiv AI
Aug 25

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

The paper introduces CLAW-4L, a benchmark of 300 pairs of English and non‑English Wikipedia biographies (French, Chinese, Azerbaijani) focused on women from non‑English contexts, complete with claim annotations and a fine‑grained claim‑pair relation corpus. It proposes a claim‑based enrichment framework that extracts claims from both biographies, aligns them to identify enrichment evidence from the non‑English version, and rewrites the English biography accordingly. Experiments demonstrate that non‑English Wikipedia biographies can improve English biography coverage, though lower‑resource settings still pose challenges.

By Yifei Song, Ziyang Chen, Emil Sayilov, Claire Gardent
arXiv Computation and Language
Sep 4

PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction

PiPMRE is a new pipeline for medical relation extraction that uses language models instead of traditional tagging schemes. The framework includes a relation generator that produces multiple relational triplets from a text and a relation filter that scores and selects the most reliable triplets. Experiments on two public datasets show that PiPMRE outperforms previous state‑of‑the‑art methods, improving recall by 5.6 points and accuracy by 4.4 points, and it also performs well in few‑shot scenarios.

By Jiaxin Duan, Fengyu Lu, Junfei Liu
Hugging Face Trending Papers
Jul 30

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks.

arXiv Computation and Language
Sep 11

CMNIE: An Information Extraction Benchmark for Chinese Military News

CMNIE is a new benchmark for extracting structured information from Chinese military news, covering event triggers, arguments, named entities, and entity relations under a unified schema. The dataset contains 13,000 manually annotated instances with 7 event types, 10 argument roles, 7 entity types, and 8 relation types. Experiments show that current supervised models, zero‑shot LLMs, and fine‑tuned LLMs struggle with relation extraction and exact span matching, highlighting the challenge of joint structured extraction in this domain.

By Yan Yu, Mengna Zhu, Zhenyu Song, Hao Yang, Haiwen Chen, Mao Wang
arXiv Computation and Language
Aug 24

Ontology-Driven Structural Regularization for Document-Level Relation Extraction

The paper introduces an ontology-driven framework to measure and enforce structural consistency in document-level relation extraction (DocRE) datasets. It identifies that many distant supervision resources, such as DocRED, contain structural noise from violations of ontology constraints and logical contradictions, which negatively affect model predictions. By incorporating structural regularization during training, the authors demonstrate a reduction in logical contradictions and improved generalization performance.

By Laura Menotti, Stefano Marchesin, Gianmaria Silvello