arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
By Fabricio F Costa
arXiv:2607. 18867v1 Announce Type: new Abstract: Large language models leak parametric knowledge of realized outcomes into historical financial decision tasks.
By Haozhe Jia
arXiv:2609.26642v1 Announce Type: cross
Abstract: Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a...
By Shivam Gupta
The paper audits a commercial non‑generative System‑1 model on biosecurity‑relevant benchmarks, evaluating accuracy, calibration, error detection, selective prediction, and sensitivity to answer‑option order. It finds that the model’s accuracy varies strongly by task, is reasonably well calibrated when the vendor’s uncertainty field is interpreted correctly, and that option order can cause significant prediction changes—averaging across rotations improves accuracy. The study also shows that applying averaging only to low‑confidence items recovers most of the gain at a lower cost.
By Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
arXiv:2607. 28399v1 Announce Type: new Abstract: Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed.
By Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li
The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.
By Dipankar Sarkar