The paper audits the impact of temporal leakage on financial-news direction prediction across 49,799 articles and 16 feature-model combinations, including TF‑IDF, MiniLM, FinBERT, and fine‑tuned RoBERTa‑large / DeBERTa‑v3‑large, as well as zero/few‑shot and LoRA probes of Llama‑3 and Qwen2.5. Random train‑test splits inflate MCC scores by 1.1× to 6.5×, with larger models and richer features showing greater gains, while end‑to‑end FinBERT fine‑tuning actually increases the gap. Only the mergers and acquisitions (M&A) category shows a positive locked‑test signal under near‑temporal chronological evaluation, with the signal localized to 2024‑2025 European‑tilted M&A semantics and not transferring to a 2009‑2020 U.S. corpus.
By Chenhao Xue, Raslen Guesmi, Siwei Feng, Yucheng Gong, Jacob Xavier Sundram, Jordan Pang, Lan Wang, Julian Kaljuvee
arXiv:2608. 08126v1 Announce Type: new Abstract: Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable.
By Gregorius Reynaldi Pratama, Kuo-Kun Tseng
arXiv:2605.24564v2 Announce Type: replace
Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable when their pre-training data include the evaluated events. An L...
By Weixian Waylon Li, Mengyu Wang, Tiejun Ma
The paper investigates how benchmark contamination—leakage of test items into training data—affects large language model (LLM) leaderboards. By comparing original test items with semantically equivalent paraphrases, the authors measure contamination as a violation of anchor-item invariance and find that it inflates absolute scores but rarely changes model rankings. Across 47 public models and 74 finetuned models on four benchmarks, the rank correlation between standard and paraphrase-controlled leaderboards is 0.997, with only a handful of cases showing differential contamination that could alter rankings.
By Xingyao Xiao (Stanford University), Yihong Cheng (City University of Macau)
arXiv:2609.14976v1 Announce Type: new
Abstract: Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage,...
By Jianhua Jiang, Dongbo Yuan, Weihua Li
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious...