Judgement in the Age of Jev: From Evaluation Scarcity to Evaluation Abundance
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.21841v1 Announce Type: new Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
arXiv:2607. 21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists.
arXiv:2609.14500v1 Announce Type: new Abstract: AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a hi...
arXiv:2607. 26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step.
arXiv:2608. 00151v2 Announce Type: replace-cross Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, efficiency, productivity, and financial return.
arXiv:2608. 01432v2 Announce Type: replace Abstract: Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin established theories of economic value.