arXiv AI By Ahmad Shahi, Mamehgol Yousefi, Brendon J. Woodford, Farhaan Mirza, Tapabrata Chakraborti

Conformal Adversarial Generative Ensemble

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Sep 17

Every Fixed Metric Has a Blind Spot: A Learned Atmospheric Critic for Scoring Forecast Realism

The paper introduces a learned atmospheric critic that discriminates between real weather data and model outputs to produce a realism score. Unlike fixed metrics, the discriminator adapts to the specific failure modes of a given model, effectively detecting various synthetic corruptions in ERA5 data. Experiments show the learned critic outperforms existing metrics and reveals that realism decreases with longer forecast lead times, favoring numerical over machine‑learning models.

By Younes Elberkennou, Dmitri Demler, Thierry Meier, Luca Rispoli, Fanny Lehmann, Joel Oskarsson
arXiv AI
Jun 30

Diversity is the Strength of the AI Crowd

arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.

By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day
Hugging Face Trending Papers
Jun 29

Diversity is the Strength of the AI Crowd

Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy?

Hugging Face Trending Papers
Jun 22

Selective Time Series Forecasting via Metalearning

Deep learning methods have achieved state-of-the-art in time series forecasting, yet their accuracy varies considerably across samples, as some instances remain inherently difficult to predict. Reject option mechanisms, which allow models to abstain from high-risk predictions, are well established in classification and regression but underexplored in forecasting.