arXiv:2609.15087v1 Announce Type: cross
Abstract: Most time series forecasting benchmarks remain numerical-centric and provide limited support for evaluating contextual information that shapes real-w...
By Peng Chen, Zhihao Zhuang, Hongzhou Chen, Junhao Huang, Aiping Yang, Mengsen Wu, Yiding Liu, Xilin Dai, Zewei Dong
arXiv:2607. 06973v1 Announce Type: new Abstract: We introduce a new context-enriched, multimodal time series forecasting benchmark, TimesX.
By Haoxin Liu, Yichen Zhou, Rajat Sen, B. Aditya Prakash, Abhimanyu Das
The paper introduces DualEvasion, a benchmark that evaluates evasion detection in earnings call Q&A using both textual transcripts and vocal cues. It contains 505 annotated question‑answer pairs from 60 calls, each labeled for textual evasion (direct vs. evasive) and speaker confidence (confident vs. unconfident). Experiments show that current multimodal models struggle to detect vocal confidence, especially in unconfident responses, and that providing speaker‑level references only modestly improves performance, leaving a significant gap compared to humans.
By Mirae Kim, Seonghun Jeong, Youngjun Kwak
arXiv:2606. 24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text.
By Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang
The paper introduces a synthetic benchmark for multimodal time‑series forecasting that evaluates how well text annotations contribute to predictions. By generating controlled signals with semantically correct, incorrect, and irrelevant annotations, the authors can precisely measure the true information content. Six mutual‑information estimators (KSG, MINE, InfoNCE, CCA, PID, and V‑information) are tested, all correctly ranking useful annotations and enabling annotation auditing without model training. The benchmark also highlights each estimator’s limitations and validates findings on seven real datasets, providing practical guidelines for metric implementation.
By Emma Andrews, Gianmarco Mengaldo
arXiv:2509.24789v5 Announce Type: replace
Abstract: The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress...
By Zhijian Xu, Wanxu Cai, Xilin Dai, Zhaorong Deng, Qiang Xu
arXiv:2609.13893v1 Announce Type: new
Abstract: Earnings conference calls are a primary channel through which managers disclose information under analyst scrutiny. Prior work has linked vocal and lex...
By Huizhong Chen, Huan Zhang
We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.
arXiv:2607. 23813v1 Announce Type: cross Abstract: We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions.
By Denglin Jiang, Haoran Zhou, Anshul Wadhawan, Brendan Fahy, Vinay Ramesh, David Weisberg, Dmitriy Derkachevskiy, Helen Sheehan, Srivas Prasad, Michele Franceschini
The paper presents the CDSP (context-conditional deliberation signal pipeline), which transforms investment committee meeting transcripts into structured predictive features. CDSP segments transcripts into topical chunks, assigns asset‑class context labels via a large language model, maps financial keywords to a taxonomy, and adds sentiment polarity and mention frequency features. Using these engineered features on 48 monthly meetings, the best model—combining sentence embeddings with CDSP features—achieves 73% accuracy and a 0.73 F1 score, outperforming a simple stock‑choice baseline, though the improvement is not statistically significant.
By Vivek Batra, Kristin Chen, Sanjiv Das, Samuel Judge, Harshad Khadilkar, Sukrit Mittal, Amir Nasrollahzadeh, Daniel Ostrov, Jacob Sisk
arXiv:2505.17613v2 Announce Type: replace
Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...
By Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu
The paper investigates whether multimodal time‑series forecasting models actually use the semantic content of accompanying text. By systematically perturbing the text—replacing it with empty, constant, shuffled, or cross‑domain sentences—the authors find that mean squared error changes by less than 0.5 % across several architectures, indicating that text does not drive performance gains. They also show that removing a co‑shipped numeric column restores the reported improvements, suggesting that the models rely on other signals rather than textual semantics.
By Karthik Sridhar, Atharva Gupta, Nishant Pradhan, Murari Mandal, Dhruv Kumar, Saurabh Deshpande