A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
arXiv:2608. 05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.
By Jin Liu, Steffen Thoma, Achim Rettinger
Large Language Models (LLMs) generate fluent long-form text, however, often add unsupported factual claims. Existing verification techniques improve factuality by grounding generation in external evidence.
arXiv:2607. 06940v1 Announce Type: cross Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality.
By Yiming Gai, Junde Lu, Xuefei Huang
arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).
By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv:2606. 18922v1 Announce Type: cross Abstract: Figurative language and negation are two areas that challenge current language models, however, both are widely used throughout written and spoken language.
By Jasmine Owers, Edwin Simpson, Martha Lewis