arXiv AI

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

TSQueryBench is a synthetic benchmark comprising 500 time‑series instances and 10 query types, each paired with correct, partially correct, and incorrect natural‑language explanations. The study evaluates six large language models on explanation generation, ranking, scoring, and anomaly detection, revealing that models often fail to generate numerically correct explanations yet can reliably identify or score correct ones. These findings suggest that rubric‑guided LLM evaluation is more dependable than generation for numerically grounded time‑series reasoning.

arXiv Computation and Language
Aug 21

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

arXiv:2608. 20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence.

By Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave, Xiao Gu, David A. Clifton
arXiv Machine Learning
Sep 17

WaveTLM: Reliable Time-Series Language Modeling through Task Compilation

WaveTLM introduces a unified compiler‑executor framework that transforms natural‑language requests into typed task states and reliable numerical or structured outputs for time‑series tasks. The authors present ExecTS‑QA, a benchmark covering forecasting, imputation, classification, anomaly detection, and waveform analysis, and show that a single WaveTLM checkpoint achieves 99.40% contract‑valid coverage versus 37.83% for a string‑first baseline while maintaining balanced predictive performance. Additional evaluations on SciTS, TSQA, IRTS‑ToolBench, and ARFBench demonstrate the model’s transferability.

By Jiahui Chen, Bingke Zhu, Hongyu Pan, Yingying Chen
arXiv AI
Jun 8

TSAQA: Time Series Analysis Question And Answering Benchmark

arXiv:2601. 23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science.

By Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, Hanghang Tong
arXiv AI
Jul 15

Scaling Point-in-Time Language Models

arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.

By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang