FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.
arXiv:2510.18368v2 Announce Type: replace Abstract: We present $\textbf{Korean SimpleQA (KoSimpleQA)}$, a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cu...
arXiv:2608. 05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.
arXiv:2605. 26937v2 Announce Type: replace-cross Abstract: Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood.
AEScorer is an agentic evidence‑grounded framework designed for graded factuality verification, addressing the limitation of binary judgments in current methods. It operates in two stages: first, it gathers and refines external evidence through agentic search; second, it predicts a scalar factuality score to capture nuanced differences in correctness. The authors also introduce GradedVeriBench, a benchmark covering general and multi‑hop question answering, and demonstrate that AEScorer outperforms existing approaches on this new benchmark.
arXiv:2606. 03883v1 Announce Type: new Abstract: Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count.
arXiv:2606. 10460v1 Announce Type: cross Abstract: Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved.
arXiv:2609.23178v1 Announce Type: new Abstract: Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its response...
The paper introduces KBevo, a co‑evolving framework that simultaneously builds a structured knowledge base and performs reasoning over it for knowledge‑intensive question answering. By optimizing both components end‑to‑end with QA outcome rewards, the system improves the quality and connectivity of the knowledge base, leading to higher answer reachability and better compositional factual reasoning. Compared to standard retrieval baselines, KBevo offers greater controllability and improved factual accuracy.
arXiv:2504. 07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness.
arXiv:2606. 27047v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge.
arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.