arXiv AI By Fabricio F Costa

Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence

Read the original on arXiv AI →

arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 24

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed
arXiv AI
Sep 11

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace is a public dataset of 558 complete AI scientific agent trajectories that records the reasoning process—thoughts, tool calls, observations, errors, revision triggers, and confidence—across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis. The dataset includes seven models (three frontier models and four open‑weight models) and 60 live‑retrieval variants, providing a balanced view of performance and error patterns. Pilot analysis shows that process traces reveal behavioral differences invisible to output‑only evaluation, such as differing error rates and types among frontier models.

By Aayam Bansal, Keertan Balaji