OpenAI Blog

Introducing SimpleQA

Read the original on OpenAI Blog →

A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

arXiv Computation and Language
Aug 28

AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification

AEScorer is an agentic evidence‑grounded framework designed for graded factuality verification, addressing the limitation of binary judgments in current methods. It operates in two stages: first, it gathers and refines external evidence through agentic search; second, it predicts a scalar factuality score to capture nuanced differences in correctness. The authors also introduce GradedVeriBench, a benchmark covering general and multi‑hop question answering, and demonstrate that AEScorer outperforms existing approaches on this new benchmark.

By Hui Huang, Muyun Yang, Yuki Arase