arXiv AI

Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks

arXiv:2501. 06286v2 Announce Type: replace-cross Abstract: Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models remain vulnerable to irrelevant context.

arXiv Computation and Language
Aug 31

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

XHotpotQA is a new benchmark for cross‑lingual knowledge composition in multi‑hop question answering. It presents each instance as an evidence‑dependency graph with explicit language assignments for the question, bridge evidence, answer‑bearing evidence, and distractors, and includes 15,661 training and 7,405 validation examples with sentence‑level support supervision. The dataset reveals significant performance drops when evidence spans language boundaries, providing a diagnostic tool for systems that must integrate evidence across languages.

By Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
arXiv Computation and Language
Sep 3

Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA

The paper introduces Evidence Sufficiency Boundary Training, a framework that teaches models to abstain from answering until the supplied evidence is fully sufficient, and to remain stable when additional redundant evidence is added. By constructing ordered evidence chains from datasets such as HotpotQA, 2WikiMultiHopQA, and MuSiQue, the method applies level supervision, a boundary flip margin, post‑boundary stability, and answer recall protection. Experiments with Qwen2.5‑3B‑Instruct and LoRA adaptation show improved boundary localization (flip accuracy 0.807 vs 0.781) and a lower unsupported‑answer rate (0.095 vs 0.101) while maintaining competitive raw QA F1.

By Haruto Sato, Yuki Tanaka, Ren Nakamura, Aoi Kobayashi, Mei Ito
arXiv Computation and Language
Sep 14

EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development

The paper introduces EAR, an Entity‑Aware Partitioning approach that improves retrieval‑augmented generation for multiple‑choice question answering by extracting normalized surface anchors from questions, answers, and the corpus. EAR retrieves local windows around matching anchors and can attach a larger parent passage via an extractive summary, reducing retrieved words by 37.5‑40.2% compared to fixed‑size chunks. Experiments on a cleaned MMLU‑style subset with Mistral, Gemma, and DeepSeek show modest accuracy changes, none statistically significant, highlighting EAR’s methodological contribution of compact, inspectable retrieval units.

By Cenab Batu Bora, Oylum Alatl{\i}, Sebnem Bora, Oguz Dikenelli
arXiv Computation and Language
Aug 27

ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives

ClueWeaver is a dual-agent framework designed to enable compact, locally deployable language models to answer questions about long literary narratives. The Finder agent retrieves passages that contain answer-critical clues, while the Interpreter agent derives the answer from those passages, generates rationales with paragraph-ID citations, and performs self-calibration for high-risk questions. Both agents are trained with reward-guided reinforcement learning to prioritize evidence retention, correctness, grounding, and concise explanations, resulting in improved performance and inspectability over end-to-end prompting.

By Jihao Zhu, Zhiwei Yang, Wenxiao Zhang, Junqian Zhao, Qi You, Fangqi Wang, Zheyuan Deng, Hanzhe Yang, Yu Liu, Jin B. Hong
arXiv AI
Jul 1

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

arXiv:2606. 32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them from the same text, and the resulting pairs are used to fine-tune, distill, or compress knowledge into another model.

By Ekaterina Alimaskina, Denis Shveykin, Gleb Molodtsov, Igor Shalygin, Alexey Kadeishvili, Aleksandr Beznosikov
arXiv AI
6d ago

Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding

The paper introduces Highlight-Then-Summarize (H2S), a two-step approach that first highlights question-relevant evidence in long documents and then condenses it into a compact, question-conditioned summary before generating an answer. The authors built the H2S-Dataset with 6,647 examples spanning 11 benchmark families, and developed H2S-RL to reward evidence selection and summary construction. Evaluated on the H2S-Bench suite, the H2S-14B model outperforms larger open-source models, achieving the highest overall score and maintaining strong performance even with a reduced output budget.

By Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc), Gaofeng Lu (Baidu Inc), Zhenyu Jiao (Baidu Inc), Xing Yuan (Baidu Inc), Dai Dai (Baidu Inc), Tong Mo (Peking University), Long Zeng (Tsinghua University)