Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.
arXiv:2510.18368v2 Announce Type: replace
Abstract: We present $\textbf{Korean SimpleQA (KoSimpleQA)}$, a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cu...
By Donghyeon Ko, Kyubyung Chae, Yeguk Jin, Byungwook Lee, Chansong Jo, Sookyo In, Jaehong Lee, Taesup Kim, Donghyun Kwak
arXiv:2608. 05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.
By Jin Liu, Steffen Thoma, Achim Rettinger
arXiv:2605. 26937v2 Announce Type: replace-cross Abstract: Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood.
By Luca Giordano, Simon Razniewski
AEScorer is an agentic evidence‑grounded framework designed for graded factuality verification, addressing the limitation of binary judgments in current methods. It operates in two stages: first, it gathers and refines external evidence through agentic search; second, it predicts a scalar factuality score to capture nuanced differences in correctness. The authors also introduce GradedVeriBench, a benchmark covering general and multi‑hop question answering, and demonstrate that AEScorer outperforms existing approaches on this new benchmark.
By Hui Huang, Muyun Yang, Yuki Arase
arXiv:2606. 03883v1 Announce Type: new Abstract: Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Fabian Farestam, Roger Wattenhofer