Reducing Hallucinations in LLM-based Scientific Literature Analysis Using Peer Context Outlier Detection
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, and that HalluPeer-defined hallucination patterns occur in real peer reviews.
HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that current detectors struggle to distinguish hallucinations from legitimate critique, and real peer reviews contain HalluPeer-defined hallucination patterns, underscoring the need for source-aware verification.
The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.
arXiv:2608.18082v2 Announce Type: replace Abstract: Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels...
arXiv:2609.15106v1 Announce Type: new Abstract: Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a laten...