arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
By Fabricio F Costa
arXiv:2607. 18867v1 Announce Type: new Abstract: Large language models leak parametric knowledge of realized outcomes into historical financial decision tasks.
By Haozhe Jia
arXiv:2609.26642v1 Announce Type: cross
Abstract: Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a...
By Shivam Gupta
The paper audits a commercial non‑generative System‑1 model on biosecurity‑relevant benchmarks, evaluating accuracy, calibration, error detection, selective prediction, and sensitivity to answer‑option order. It finds that the model’s accuracy varies strongly by task, is reasonably well calibrated when the vendor’s uncertainty field is interpreted correctly, and that option order can cause significant prediction changes—averaging across rotations improves accuracy. The study also shows that applying averaging only to low‑confidence items recovers most of the gain at a lower cost.
By Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
arXiv:2607. 28399v1 Announce Type: new Abstract: Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed.
By Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li
The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.
By Dipankar Sarkar
arXiv:2606. 10794v3 Announce Type: replace Abstract: Existing black-box LLM provenance methods achieve comparability by querying every candidate model with the same diagnostic prompts.
By Jiaxu Liu, Sunnan Mu, Dong Huang, Liuyin Wang, Jing Shao, Jie Zhang
The paper introduces VINTAGE-TS, a revision‑aware time‑series foundation model that separates observation time from information‑availability time. It predicts both the next period’s first‑published value and the value available after a fixed delay, maintaining a joint distribution to capture their dependence and uncertainty. The authors provide a detailed evaluation protocol, software tools for validity‑interval reconstruction and delayed‑label filtering, and a synthetic demonstration with a 25‑configuration sensitivity suite to illustrate performance variability and the impact of hindsight contamination.
By Taimoor Ahmad
The paper introduces instruction duplication, a simple inference‑time control that repeats the procedural instruction without retraining or decoding changes. Across seven instruction‑tuned models and 16,800 scheduled generations, duplicating the instruction improves deterministic All‑8 diagnostic‑response success from 90.22% to 93.17% and reduces failures by 30.2%. In downstream Answer Engineering scenarios, duplication further boosts success rates, demonstrating its practical impact on systems that rely on the generated trajectory.
By Victor Lavrenko (PeaceTech VC, Israel)
The paper argues that zero‑shot time‑series forecasting should be treated as an evidence‑access claim rather than merely a no‑parameter‑update condition. It introduces a source‑first taxonomy that distinguishes three evidence sources—frozen LLM prior reuse, parametric time‑series pretraining, and retrieval‑augmented external memory—from the architectures that implement them. The authors further outline four audit questions—task interface, forecast object and scoring, prediction‑time context, and resource budget—to make zero‑shot leaderboards transparent and comparable.
By Delun Kong, Wanyun Ling, Chenxi Liu, Ziyue Li
The paper investigates when forecasting agents should employ different behaviors—retrieval, reasoning, deferring to market priors, or using historical analogs—on binary forecasting tasks. It finds that the optimal mechanism depends on the data source, with structured analogs excelling for some processes and market or conservative baselines for others. The authors propose ReliabilityRoute, a rule‑based system that steers agent behavior using reliability features, achieving competitive performance across multiple LLM versions while highlighting that more reasoning is not always better.
By Yufeng Wang
Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast.