arXiv AI

Judgement in the Age of Jev: From Evaluation Scarcity to Evaluation Abundance

arXiv AI
Aug 5

A New Theory of Value for Post-AGI Economics

arXiv:2608. 01432v2 Announce Type: replace Abstract: Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin established theories of economic value.

By Keyun Ruan
arXiv AI
Jun 6

SAGE: Scalable AI Governance & Evaluation

arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.

By Benjamin Le, Xueying Lu, Nick Stern, Wenqiong Liu, Igor Lapchuk, Xiang Li, Baofen Zheng, Kevin Rosenberg, Jiewen Huang, Zhe Zhang, Abraham Cabangbang, Satej Milind Wagle, Jianqiang Shen, Raghavan Muthuregunathan, Abhinav Gupta, Mathew Teoh, Andrew Kirk, Thomas Kwan, Jingwei Wu, Wenjing Zhang
arXiv AI
Sep 24

Spread and Scale: What Determines Whether Test-Time Budget Allocation Pays

The paper investigates when reallocating a fixed test‑time budget toward harder instances improves solution quality for neural combinatorial optimization solvers. Through pre‑registered experiments on three solvers and two hard‑workload constructions for the traveling salesman problem, it finds that the key deciding factor is the variation in instance difficulty within a workload, not the average difficulty. A budget‑aware policy that first spends part of the budget to gauge instance difficulty recovers most of the potential improvement, though not all, when the cost of this information is included.

By Jinhyung Bae