arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2607. 21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists.
By Chen Zhu, Xiaolu Wang, Weilong Zhang
arXiv:2609.14500v1 Announce Type: new
Abstract: AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a hi...
By Seyed Morteza Emadi
arXiv:2607. 26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step.
By Brett Reynolds
arXiv:2608. 00151v2 Announce Type: replace-cross Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, efficiency, productivity, and financial return.
By Keyun Ruan, Jonathan D. Teubner, John M. Bremen
arXiv:2608. 01432v2 Announce Type: replace Abstract: Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin established theories of economic value.
By Keyun Ruan
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own?
arXiv:2607. 15164v1 Announce Type: new Abstract: Artificial intelligence is transforming scientific research - not merely as a more powerful instrument, but as an autonomous participant in the research cycle itself.
By Emmanuel Jeannot
arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.
By Benjamin Le, Xueying Lu, Nick Stern, Wenqiong Liu, Igor Lapchuk, Xiang Li, Baofen Zheng, Kevin Rosenberg, Jiewen Huang, Zhe Zhang, Abraham Cabangbang, Satej Milind Wagle, Jianqiang Shen, Raghavan Muthuregunathan, Abhinav Gupta, Mathew Teoh, Andrew Kirk, Thomas Kwan, Jingwei Wu, Wenjing Zhang
arXiv:2608.29843v1 Announce Type: cross
Abstract: Posted prices for AI inference have fallen steadily since 2024, yet the measured speed of that fall depends almost entirely on the method of measurem...
By Louis Yiven Zhu
arXiv:2608. 08882v1 Announce Type: cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present.
By Christoph Trattner
The paper investigates when reallocating a fixed test‑time budget toward harder instances improves solution quality for neural combinatorial optimization solvers. Through pre‑registered experiments on three solvers and two hard‑workload constructions for the traveling salesman problem, it finds that the key deciding factor is the variation in instance difficulty within a workload, not the average difficulty. A budget‑aware policy that first spends part of the budget to gauge instance difficulty recovers most of the potential improvement, though not all, when the cost of this information is included.
By Jinhyung Bae