MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.15087v1 Announce Type: cross Abstract: Most time series forecasting benchmarks remain numerical-centric and provide limited support for evaluating contextual information that shapes real-w...
arXiv:2607. 06973v1 Announce Type: new Abstract: We introduce a new context-enriched, multimodal time series forecasting benchmark, TimesX.
The paper introduces DualEvasion, a benchmark that evaluates evasion detection in earnings call Q&A using both textual transcripts and vocal cues. It contains 505 annotated question‑answer pairs from 60 calls, each labeled for textual evasion (direct vs. evasive) and speaker confidence (confident vs. unconfident). Experiments show that current multimodal models struggle to detect vocal confidence, especially in unconfident responses, and that providing speaker‑level references only modestly improves performance, leaving a significant gap compared to humans.
arXiv:2606. 24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text.
The paper introduces a synthetic benchmark for multimodal time‑series forecasting that evaluates how well text annotations contribute to predictions. By generating controlled signals with semantically correct, incorrect, and irrelevant annotations, the authors can precisely measure the true information content. Six mutual‑information estimators (KSG, MINE, InfoNCE, CCA, PID, and V‑information) are tested, all correctly ranking useful annotations and enabling annotation auditing without model training. The benchmark also highlights each estimator’s limitations and validates findings on seven real datasets, providing practical guidelines for metric implementation.
arXiv:2509.24789v5 Announce Type: replace Abstract: The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress...