arXiv:2606. 05724v1 Announce Type: cross Abstract: Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing character states, social relations, causal triggers, temporal position, and later consequences.
By Qiuyu Tian, Fengyi Chen, Yiding Li, Youyong Kong, Fan Guo, Yuyao Li, Jinjing Shen, Zhijing Xie, Yiyun Luo, Xin Zhang, Yingce Xia, Zequn Liu
arXiv:2605.27156v2 Announce Type: replace-cross
Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, particularly for long-tail do...
By Ruikang Zhang, Zhanni Chen, Yiqiao Cai, Qi Su
arXiv:2609.38799v1 Announce Type: new
Abstract: Understanding multi-perspective alternative narratives requires identifying how their information agrees, conflicts, or differs across sources. Existin...
By Eftekhar Hossain, Santu Karmaker
NS-ST-GraphRAG is a neuro‑symbolic spatio‑temporal GraphRAG framework designed to process long‑form literary narratives by integrating ontology‑guided extraction, deterministic constraint checking, dual temporal coordinates, spatial scene attributes, and dynamic sub‑graph retrieval. It selects the appropriate graph state based on the temporal and spatial scope of a query, grounding generated answers in traceable evidence. The authors also introduce Red‑Chamber‑QA, an open multi‑hop question‑answering benchmark for classical Chinese literature, and report that NS‑ST‑GraphRAG outperforms a frozen‑window baseline and a closed‑book model on a held‑out 120‑question split.
By Zheng Kui Lin
arXiv:2607. 09328v2 Announce Type: replace-cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages.
By Zixin Chen, Peng Liu, Haobo Li, Rui Sheng, Jianhong Tu, Xiaodong Deng, Fei Huang, Kashun Shum, Dayiheng Liu, Huamin Qu
arXiv:2607. 09328v1 Announce Type: cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages.
By Zixin Chen, Peng Liu, Haobo Li, Rui Sheng, Jianhong Tu, Xiaodong Deng, Fei Huang, Kashun Shum, Dayiheng Liu, Huamin Qu
This study tackles the challenge of identifying essay‑scale republication and reuse from fragmented text evidence, focusing on David Hume essays in eighteenth‑century books and newspapers. It compares a staged rule‑based workflow, baseline decision‑tree and LLM approaches, and automated rule adaptation, finding that pair‑level feature aggregation achieves high F1 scores and that the final workflow offers the best precision‑recall balance. Manual audits confirm all predicted positives as genuine republications, demonstrating the method’s effectiveness in producing compact, auditable candidate sets for historical analysis.
By Ke Shu, Kira Hinderks, Eetu M\"akel\"a, Mikko Tolonen
arXiv:2607. 05577v1 Announce Type: new Abstract: Long-form fiction writers need memory that answers multi-hop questions about evolving story state: who knows a secret and when they learned it, whether an event preceded the narration that revealed it, whether a setup paid off, and how a relationship shifted.
By Mohammad Saifullah, Thomas Kornmaier, Taaha Kazi, Vasu Sharma, Aditya Sanjiv Kanade, Aanand Kumar Yadav
arXiv:2609.00241v1 Announce Type: new
Abstract: Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particular...
By Meng Zhou, Wenhao You, Wei Yuan
arXiv:2609.38406v1 Announce Type: new
Abstract: Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct mi...
By Eftekhar Hossain, John Salvador, Santu Karmaker
The paper introduces a three-tier evaluation framework—recovery, mining, and discovery—for unsupervised narrative label generation in disinformation datasets. It compares clustering-based and graph-community-based pipelines across seven datasets, finding that clustering can underrepresent prominent topics while graph methods produce many singletons that human annotators recognize as valid narratives. The authors release human-validated narrative candidate labels for the Climate Obstruction and PolyNarrative datasets to aid taxonomy development and dataset expansion.
By Max Upravitelev, Veronika Solopova, Jing Yang, Charlott Jakob, Alexandra Tsiakalou, Neda Foroutan, Vera Schmitt
The paper presents a detailed examination of narrative elements—agency, setting, and events—within the Dolma web-scale pretraining corpus. Using a framework of 11 interpretable dimensions, the authors hand‑annotated 400 passages, expanded this to a 25,000‑passage LLM‑labeled dataset, and trained NarraBERT models to predict narrative features across 13 million passages, producing the NarraDolma dataset. The study reveals that narrative structure is measurable at scale and that narrative qualities vary unevenly across different data sources, topics, and formats, highlighting gaps in current data curation practices.
By Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak