arXiv:2608. 08882v1 Announce Type: cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present.
By Christoph Trattner
arXiv:2606. 31273v1 Announce Type: new Abstract: AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them.
By Hongmin Li
arXiv:2607. 26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step.
By Brett Reynolds
arXiv:2609.15624v1 Announce Type: cross
Abstract: Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We...
By Daniele Veri'
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.
By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan