arXiv AI

Using LLMs for Explainable, Data-Driven Insight Generation from Time Series

arXiv:2607. 18271v1 Announce Type: new Abstract: Time series forecasts are widely used in decision-critical domains, where they are rarely consumed without accompanying explanations.

arXiv AI
Aug 19

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

TSQueryBench is a synthetic benchmark comprising 500 time‑series instances and 10 query types, each paired with correct, partially correct, and incorrect natural‑language explanations. The study evaluates six large language models on explanation generation, ranking, scoring, and anomaly detection, revealing that models often fail to generate numerically correct explanations yet can reliably identify or score correct ones. These findings suggest that rubric‑guided LLM evaluation is more dependable than generation for numerically grounded time‑series reasoning.

By Preetham Sivalingam, Murari Mandal, Dhruv Kumar, Saurabh Deshpande
arXiv AI
3d ago

Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

The paper introduces a framework that uses large language models (LLMs) to generate natural‑language narratives explaining cross‑sectional stock return predictions. It combines temporal Shapley additive explanations (SHAP) from an XGBoost model with historical regime analogs to provide context. A controlled study shows that progressively externalizing numerical and relational reasoning improves evidence faithfulness and accuracy, while historical analogs boost human‑rated usefulness.

By Sujung Kim, Seung Hwan Cho, Sangjin Park, Young-Min Kim
arXiv Machine Learning
Aug 27

NVExplain: Explaining Time Series Forecasting with Latent Trajectory Analysis and Structure-Preserving Surrogates

NVExplain is a model‑agnostic framework that explains time‑series forecasting by attributing each forecast horizon to temporally relevant historical lags. It models forecasting as a latent trajectory, introduces semantic flow to track information evolution, and aggregates this into a lag‑horizon attribution matrix. The method also generates structure‑preserving perturbations and fits sparse local surrogates to produce human‑readable, temporally coherent explanations, and demonstrates competitive faithfulness and stability across benchmark datasets.

By Muyan Anna Li, Manikandan Ravikiran, Aditi Gautam