arXiv AI

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

arXiv:2608. 11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists.

arXiv AI
Jul 16

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

arXiv:2601. 02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctuates significantly depending on how information is distributed in real-world corpora.

By Amirali Ebrahimzadeh, Seyyed M. Salili