arXiv AI By Preetham Sivalingam, Murari Mandal, Dhruv Kumar, Saurabh Deshpande

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

Read the original on arXiv AI →

TSQueryBench is a synthetic benchmark comprising 500 time‑series instances and 10 query types, each paired with correct, partially correct, and incorrect natural‑language explanations. The study evaluates six large language models on explanation generation, ranking, scoring, and anomaly detection, revealing that models often fail to generate numerically correct explanations yet can reliably identify or score correct ones. These findings suggest that rubric‑guided LLM evaluation is more dependable than generation for numerically grounded time‑series reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 21

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

arXiv:2608. 20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence.

By Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave, Xiao Gu, David A. Clifton
arXiv Machine Learning
Sep 17

WaveTLM: Reliable Time-Series Language Modeling through Task Compilation

WaveTLM introduces a unified compiler‑executor framework that transforms natural‑language requests into typed task states and reliable numerical or structured outputs for time‑series tasks. The authors present ExecTS‑QA, a benchmark covering forecasting, imputation, classification, anomaly detection, and waveform analysis, and show that a single WaveTLM checkpoint achieves 99.40% contract‑valid coverage versus 37.83% for a string‑first baseline while maintaining balanced predictive performance. Additional evaluations on SciTS, TSQA, IRTS‑ToolBench, and ARFBench demonstrate the model’s transferability.

By Jiahui Chen, Bingke Zhu, Hongyu Pan, Yingying Chen
arXiv AI
Jun 8

TSAQA: Time Series Analysis Question And Answering Benchmark

arXiv:2601. 23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science.

By Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, Hanghang Tong