Hugging Face Trending Papers

STeMP: Spatio-Temporal Modelling Protocol

Spatio-temporal machine-learning modelling is an important tool in environmental research. However, machine-learning models are highly sensitive to both the characteristics of the training data, such as its distribution, and methodological choices, including the cross-validation strategy.

arXiv Machine Learning
Jul 24

STeMP: Spatio-Temporal Modelling Protocol

arXiv:2607. 20592v1 Announce Type: new Abstract: Spatio-temporal machine-learning modelling is an important tool in environmental research.

By Jan Linnenbrink, Jakub Nowosad, Marvin Ludwig, Anna Frederike Jablotschkin, Fabian Schumacher, Teja Kattenborn, Hanna Meyer
arXiv Machine Learning
Jul 2

OpFML: Pipeline for ML-based Operational Inference

arXiv:2601. 11046v2 Announce Type: replace Abstract: Machine learning models for climate and Earth science are becoming increasingly capable, yet model deployment into operational use remains a largely unaddressed challenge: general-purpose model-serving tools, such as MLflow and KServe, assume input data availability at the inference node, while data acquisition, failure handling, and preprocessing are trusted to a separate workflow.

By Shahbaz Alvi, Giusy Fedele, Gabriele Accarino, Italo Epicoco, Ilenia Manco, Pasquale Schiano
arXiv Machine Learning
Aug 27

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

The Planetary Prediction Engine (PPE) is an autonomous AI system that transforms natural-language queries into end-to-end geospatial predictions. It automatically retrieves and fuses multimodal datasets from open-web and Earth observation sources, incorporates foundation model embeddings, and searches task‑specific model families with overfitting safeguards. Across multiple domains, PPE outperforms state‑of‑the‑art baselines, improving regression metrics for CDC health indicators, FEMA risk indices, and the Social Vulnerability Index, doubling accuracy for Nigerian food security indicators, and achieving higher recall in Ebola outbreak nowcasting.

By Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun, Hamed Sadeghi, Dav M. Ebengo, Mbulayi Onesime, Rouslan Solomakhin, John Wamburu, William Ogallo, Aisha Walcott-Bryant, Sanxing Chen, Arbaaz Muslim, Yael Mayer, Ronald Ho, Roy Lee, Ruth Alcantara, Abdoulaye Diack, Monica Bharel, Lambert Rosique, Jeremy Amez-Droz, Christopher Haire, James Manyika, Yossi Matias, Niv Efron, Gautam Prasad, Shravya Shetty
Hugging Face Trending Papers
Sep 3

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench is a dynamic interactive benchmark designed to evaluate how large language model agents reconcile user instructions, internal knowledge, and real‑time environmental observations. It consists of 238 manually curated multi‑turn tasks that test world‑knowledge conflicts, input inconsistencies, and multi‑source temporal conflicts, using a user simulator, stateful tools, deterministic environment assertions, an open‑source natural‑language evaluator, and human trajectory verification. Evaluation of nine models—including DeepSeek‑V4‑Flash, GLM‑5.2, and MiniMax‑M3—reveals significant cross‑domain variation, with no model reliably handling factual correction, identity consistency, and temporal conflict resolution across all settings, and shows that missed conflicts can propagate to tool calls or synthetic protected‑data flows.