TRACE is a training‑free, agentic retrieval framework that enables accountable source discovery in historical archives, addressing challenges such as OCR degradation and genre heterogeneity. Developed for the DECIDON project on French Third Republic political discourse, it is deployed internally for 24 researchers across six institutions. On the HistoriQA‑ThirdRepublic benchmark, TRACE achieves R@10 of 0.856 and MRR of 0.653, outperforming sparse, dense, graph‑based, and other agentic RAG baselines, especially on multi‑hop and cross‑corpus questions, while costing only about $0.02 per question.
By Donghan Bian (ENC, LRE), Marie Puren (LRE, ENC), Florian Cafiero (LRE, ENC)
Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-2026, the third edition of the HIPE evaluation series.
arXiv:2609.23178v1 Announce Type: new
Abstract: Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its response...
By Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Edwin Roland, Wenyi Shang, Matthew Wilkens
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
Trustworthy Legal AI requires systems that can answer legal questions while grounding their responses in authoritative sources. However, existing Vietnamese legal benchmarks provide limited coverage o...
arXiv:2606. 10460v1 Announce Type: cross Abstract: Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved.
By Haonan Wang, Jiaxiang Liu, Yurong Liu, Austin Senna Wijaya, Tianle Zhou, Eden Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto, Grace Fan, Yusen Zhang, Juliana Freire, Eugene Wu
ViLegalExpert is a large-scale Vietnamese legal benchmark built from real citizen–lawyer consultations, comprising over 172,000 questions across 34 legal domains with professional answers and expert-verified evidence. It supports legal information retrieval, extractive QA, and abstractive QA. Experiments show that while pretrained language models perform well on QA, hybrid retrieval methods achieve the best evidence retrieval, highlighting significant challenges in grounding legal answers to authoritative sources.
By Dat Tien Nguyen, Nghia Hieu Nguyen, Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
arXiv:2606. 00029v1 Announce Type: cross Abstract: Retrieval-augmented generation systems struggle with temporal reasoning and evidence fusion when answering complex questions over historical criminal case narratives.
By Sidra Nasir, Muhammad Noman Zahid, Rizwan Ahmed Khan
arXiv:2609.14528v1 Announce Type: cross
Abstract: Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along paths in a KG to answer natural-langua...
By Eduin E. Hernandez, Luis F. Garcia, Nurassyl Askar, Sergio A. Diaz, Stefano Rini
The paper explores using large language models (LLMs) as AI respondents to convert policy documents into structured survey responses. It introduces a long-context in‑context learning pipeline that maps policy text to predefined survey categories such as policy instruments, target groups, and thematic areas, and includes a secondary LLM validation step. Evaluation on a multi‑country dataset shows high agreement (84‑95%) with human responses for structured indicators, though free‑text fields differ, indicating that hybrid human‑AI workflows can enhance policy monitoring efficiency while still requiring human oversight.
By Carolyn Cole, Matthias Deschryvere, Toqeer Ehsan, Arash Hajikhani
arXiv:2509. 21028v4 Announce Type: replace Abstract: We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs).
By Miao Li, Alexander Gurung, Irina Saparina, Mirella Lapata
arXiv:2608.21252v1 Announce Type: cross
Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationshi...
By Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han