arXiv AI By Nathan Thierry, Andre-Louis Rochet

TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training Split

Read the original on arXiv AI →

arXiv:2609. 28506v1 Announce Type: new Abstract: TW3Cast is a time-series forecasting system that reaches position 3 of 130 entries on the GIFT-Eval benchmark by mean MASE rank, as of 2026-09-14.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed
arXiv Computation and Language
Sep 10

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

MetroLLM-Bench is a 955‑case benchmark designed to evaluate language models as the policy layer of transit kiosks across six real metro systems, covering routing, fare calculation, disruptions, accessibility, and adversarial input. The benchmark includes 14 deterministic scoring components (Tier 1) and 8 semantic‑quality components (Tier 2), with a 75/25 split for training‑data generation and held‑out evaluation. Twenty‑six models from six vendors were tested, and a 4B Qwen 3.5 student fine‑tuned via PEFT outperformed GPT‑5.6 on Tier 1 and matched GPT‑5.4 on the combined score, while larger models offered no further improvement.

By Remco Hendriks (Continker)