Hugging Face Trending Papers

Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts

Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-2026, the third edition of the HIPE evaluation series.

arXiv AI
Jul 1

HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

arXiv:2606. 31325v1 Announce Type: new Abstract: We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic.

By Aur\'elien Pellet (LRE), Julien Perez (EPITA, LRE), Marie Puren (LRE, CJM)
arXiv AI
Aug 25

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

The paper introduces CLAW-4L, a benchmark of 300 pairs of English and non‑English Wikipedia biographies (French, Chinese, Azerbaijani) focused on women from non‑English contexts, complete with claim annotations and a fine‑grained claim‑pair relation corpus. It proposes a claim‑based enrichment framework that extracts claims from both biographies, aligns them to identify enrichment evidence from the non‑English version, and rewrites the English biography accordingly. Experiments demonstrate that non‑English Wikipedia biographies can improve English biography coverage, though lower‑resource settings still pose challenges.

By Yifei Song, Ziyang Chen, Emil Sayilov, Claire Gardent
arXiv AI
Sep 10

Beyond Single-Negative Preference: Multi-Negative DPO for LLM-Centric Historical Entity Linking

The paper introduces Multi-Negative Direct Preference Optimisation (MDPO), a pairwise objective that compares the correct entity with all valid rejected candidates for each mention, extending the single-negative approach used in prior work. MDPO retains the Bradley‑Terry formulation of Direct Preference Optimisation while leveraging the full candidate set through masked, length‑normalised sequence scores. Experiments on French, German, English, Swedish, and Finnish historical newspaper datasets (hipe‑2020 and newseye) show that MDPO outperforms both supervised fine‑tuning and single‑negative DPO, especially for NIL mentions, semantic ambiguity, OCR noise, and historically challenging names, and highlight candidate retrieval as a key bottleneck.

By Tien Nam Nguyen, Emanuela Boros, Ahmed Hamdi, Adam Jatowt, Micka\"el Coustaty, Antoine Doucet
Hugging Face Trending Papers
Jul 30

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks.

arXiv AI
Sep 18

TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives

TRACE is a training‑free, agentic retrieval framework that enables accountable source discovery in historical archives, addressing challenges such as OCR degradation and genre heterogeneity. Developed for the DECIDON project on French Third Republic political discourse, it is deployed internally for 24 researchers across six institutions. On the HistoriQA‑ThirdRepublic benchmark, TRACE achieves R@10 of 0.856 and MRR of 0.653, outperforming sparse, dense, graph‑based, and other agentic RAG baselines, especially on multi‑hop and cross‑corpus questions, while costing only about $0.02 per question.

By Donghan Bian (ENC, LRE), Marie Puren (LRE, ENC), Florian Cafiero (LRE, ENC)
arXiv Computation and Language
6d ago

From annotation to reasoning: Culture in language models

The paper proposes a new way to evaluate language models on cultural understanding by focusing on interpretive depth rather than just factual recall. It argues that literary interpretation, where scholars can disagree yet still assess the quality of evidence, provides a useful framework for testing how models handle cultural references, reuse, and transformation across texts. The authors suggest linking evidence-centered benchmarks, preserving scholarly disagreement, and conducting model-development experiments on literary data, with Danish literature as a starting point for broader applications.

By Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen
arXiv Computation and Language
Sep 11

CMNIE: An Information Extraction Benchmark for Chinese Military News

CMNIE is a new benchmark for extracting structured information from Chinese military news, covering event triggers, arguments, named entities, and entity relations under a unified schema. The dataset contains 13,000 manually annotated instances with 7 event types, 10 argument roles, 7 entity types, and 8 relation types. Experiments show that current supervised models, zero‑shot LLMs, and fine‑tuned LLMs struggle with relation extraction and exact span matching, highlighting the challenge of joint structured extraction in this domain.

By Yan Yu, Mengna Zhu, Zhenyu Song, Hao Yang, Haiwen Chen, Mao Wang
arXiv Computation and Language
Sep 14

Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

The paper introduces a weakly supervised framework for extracting dataset mentions from forced displacement and Fragile, Conflict, and Violence (FCV) documents. It uses a lightweight model trained on general research literature to generate candidate mentions, which are then refined by a large language model that validates or rejects them and corrects boundaries. The refined annotations are augmented with synthetic and contrastive examples to fine‑tune the model, achieving 74.1% precision and 70.5% recall on a benchmark of 1,706 passages, with higher precision (89.5%) on passages that contain dataset references.

By Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez