arXiv:2605. 04135v2 Announce Type: replace-cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do.
By David Gringras, Misha Salahshoor
arXiv:2607. 19453v1 Announce Type: cross Abstract: We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs.
By Ayoub Jadouli
arXiv:2607. 10972v1 Announce Type: new Abstract: Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop.
By Aleh Manchuliantsau
arXiv:2607. 12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets.
By Taizhen Cheung, SA Kwon
arXiv:2607. 05450v1 Announce Type: cross Abstract: This paper explores the "Granularity Paradox" in time-series forecasting, wherein finer temporal disaggregation (e.
By Hugo Moreira
arXiv:2608. 10433v2 Announce Type: replace Abstract: Temporal reports are increasingly emitted alongside numerical forecasts and are often interpreted as statements about the computation producing those forecasts.
By Qipeng Qian, Yuntao Qian
arXiv:2607. 08046v1 Announce Type: cross Abstract: Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast.
By Rapha\"el Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho
arXiv:2605. 22681v2 Announce Type: replace Abstract: AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances.
By Sean Wu, Pan Lu, Yupeng Chen, Jonathan Bragg, Yutaro Yamada, Peter Clark, David Clifton, Philip Torr, James Zou, Junchi Yu
arXiv:2606. 04342v1 Announce Type: cross Abstract: Multi-step time series forecasting (MSF) is commonly evaluated using point-wise error metrics such as mean squared error (MSE), implicitly treating the conditional mean as a sufficient target.
By Riku Green, Zahraa S. Abdallah, Telmo M Silva Filho
arXiv:2606. 02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation pipelines lack provenance aware auditability, hidden drift detection, and reproducibility oriented governance.
By Karan Sehgal, Khawar Naveed Bhatti
Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast.
arXiv:2606. 09473v1 Announce Type: cross Abstract: Probabilistic forecasters are increasingly learned, yet the baselines they are compared against are often weak or omitted.
By Valery Manokhin