HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, and that HalluPeer-defined hallucination patterns occur in real peer reviews.
By Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai
HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that current detectors struggle to distinguish hallucinations from legitimate critique, and real peer reviews contain HalluPeer-defined hallucination patterns, underscoring the need for source-aware verification.
The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
By Valentin Romanov, Monique Bax, Steven Niederer
The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.
arXiv:2608.18082v2 Announce Type: replace
Abstract: Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels...
By Ruizhi Zhang, Jinwei Chen, Xiangju Lu, He Yan, Mo Yu, Junmin Zhu, Wei Zhang
arXiv:2609.15106v1 Announce Type: new
Abstract: Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a laten...
By Xuhan Tong, Jiawei Zhang
arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.
By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
The paper presents an empirical study of factual errors in human-written text, focusing on corrections in newspaper articles to build a taxonomy of common mistakes such as kanji misconversions and unit errors. It evaluates large language models’ ability to detect these errors, finding that even advanced models like GPT‑5.4 achieve only a 52% word‑level F1 score on synthetic data, underscoring the difficulty of the task. The work highlights the gap in research on factual error detection in human writing compared to LLM hallucinations.
By Kazuma Iwamoto, Kazumasa Omura, Shotaro Ishihara
The paper examines a specific type of hallucination in large language models caused by spurious correlations—unintended, statistically prominent associations in training data such as surnames linked to nationalities. These hallucinations are confidently produced, persist regardless of model scaling or refusal fine‑tuning, and evade existing detection methods like confidence filtering and inner‑state probing. The authors use controlled synthetic experiments and evaluations on both open‑source and proprietary LLMs, including GPT‑5, to demonstrate the failure of current detection techniques and provide a theoretical explanation for why statistical biases undermine confidence‑based approaches.
By Shaowen Wang, Yiqi Dong, Ruinian Chang, Tansheng Zhu, Yuebo Sun, Kaifeng Lyu, Jian Li
arXiv:2606. 06748v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) reduces but does not eliminate hallucination in large language models.
By Jianru Shen
arXiv:2608. 10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing.
By Lena Holzwarth, Rita Gonz\'alez-M\'arquez, Dmitry Kobak
The paper introduces REASONS, a benchmark comprising 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution by large language models. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to assess the trade-off between reliability and responsiveness. Experiments on proprietary and open-source LLMs under various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but may increase abstention, while retrieval-augmented variants often maintain near-zero abstention. Human evaluation reveals a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.
By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur