The paper introduces REASONS, a benchmark comprising 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution by large language models. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to assess the trade-off between reliability and responsiveness. Experiments on proprietary and open-source LLMs under various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but may increase abstention, while retrieval-augmented variants often maintain near-zero abstention. Human evaluation reveals a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.
By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur
arXiv:2608.21376v1 Announce Type: cross
Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against...
By Yu Hou, Hal Daum\'e III, Rachel Rudinger, William Walden
The paper introduces REASONS, a benchmark of 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution under different evidence conditions. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to balance reliability and responsiveness. Experiments with proprietary and open-source LLMs across various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but increases abstention, while adversarial metadata can push hallucination rates above 85%. Human evaluation confirms a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.
By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur
arXiv:2609.01432v1 Announce Type: cross
Abstract: Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentionin...
By Yixuan Liu, Lin Chen, Zhuoqi Liu, Jianglin Lu, Dakota Murray
arXiv:2606. 28358v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in external documents, often using inline citations for verifiability.
By Ian van Dort (University of Amsterdam), Maria Heuss (University of Amsterdam)
arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.
By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu
The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
By Valentin Romanov, Monique Bax, Steven Niederer
arXiv:2608. 05157v1 Announce Type: cross Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias.
By Bulambo Mwendelwa Gloire, Prasenjit Mitra
The study examines whether large language models (LLMs) used as conversational search engines for academic literature prioritize papers based on authority signals—such as author prestige, venue, and citations—rather than content. By keeping titles and abstracts constant and manipulating authority metadata across three counterfactual conditions (original, flipped, boosted), the researchers tested eight LLMs in a single-turn, top‑1 recommendation scenario. Results reveal a substantial, directional authority bias that varies across models and is only partially mitigated by prompt-level debiasing, while also highlighting a significant say‑do gap where debiasing instructions suppress authority mentions more quickly than authority-driven flips, leading to underestimation of behavioral bias.
By Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki, Benjamin M. Ampel, Rajshekhar Sunderraman, Yi Ding
The paper introduces AttriBench, a benchmark dataset that balances author fame and demographics to study quote attribution in large language models (LLMs). Using AttriBench, the authors evaluate 11 popular LLMs and find that accurate attribution remains difficult, with significant disparities across race, gender, and intersectional groups. They also identify a new failure mode—suppression—where models omit attribution entirely, which is unevenly distributed across demographics and not reflected by standard accuracy metrics.
By Eliza Berman, Bella Chang, Daniel B. Neill, Emily Black
The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.
arXiv:2603. 26791v3 Announce Type: replace-cross Abstract: Assessing a cited paper's impact is typically done by analyzing its citation context in isolation within the citing paper.
By Hannah Collison, Benjamin Van Durme, Daniel Khashabi