TimeInteract introduces a new regime called Time-Series Interaction, enabling models to continuously perceive incoming time-series data and user intent, decide when to respond, and keep processing new observations during response generation. The system employs a dual-view streaming encoder, a response control mechanism, and a decoupled inference pipeline to avoid blocking. Evaluated on the newly created StreamTSI-34K dataset, TimeInteract outperforms existing LLMs, VLMs, and TSLMs across four interaction levels, achieving significant gains in accuracy, response triggering, and inference speed.
By Sheng Pan, Yongli Gu, Yiqing Guo, Warren Jin, Bo Du, Shirui Pan, Ming Jin
The paper proposes a new approach to window‑level monitoring of temporal processes by treating it as a reference‑based hypothesis test. Instead of relying on predefined parametric models, the method uses an empirical reference distribution derived from task‑ or domain‑specific data, combined with pretrained time‑series encoders, kernel density estimation, and conformal calibration to provide finite‑sample valid inference in a learned representation space. Classical concepts such as stationarity and cyclostationarity naturally emerge as special cases of this framework, and experiments show the method’s sensitivity to distributional changes while maintaining well‑calibrated inference under stable conditions.
By Jinmyeong Choi, Taesup Kim, Artur Dubrawski
arXiv:2603. 11756v2 Announce Type: replace Abstract: Deep generative models for anomaly detection in multivariate time-series are typically trained by maximizing observed data likelihood.
By David Baumgartner, Eliezer de Souza da Silva, I\~nigo Urteaga
arXiv:2606. 16863v1 Announce Type: new Abstract: Evaluation of spatiotemporal point process (STPP) models relies heavily on opaque real-world datasets, where latent generative structure is unknown and model failures are difficult to attribute.
By Yahya Aalaila, Sumantrak Mukherjee, Gerrit Gro{\ss}mann, Sebastian Vollmer
arXiv:2606. 28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting.
By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
arXiv:2602.01605v2 Announce Type: replace
Abstract: Time Series Foundation Models (TSFMs) leverage extensive pretraining to accurately predict unseen time series during inference, without the need fo...
By Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai, William Gilpin
arXiv:2607. 25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks.
By Xiaoyu Huang, Lulu Wang
arXiv:2604. 17616v3 Announce Type: replace Abstract: Root cause analysis (RCA) for time-series anomaly detection is critical for the reliable operation of complex real-world systems.
By Shashank Mishra, Karan Patil, Cedric Schockaert, Didier Stricker, Jason Rambach
arXiv:2501. 04339v2 Announce Type: replace-cross Abstract: We introduce the Deep Convolutional Interpreter for Time Series (DCIts), a deep-learning architecture for nonlinear multivariate time series that provides sample-specific, locally interpretable descriptions of the underlying interaction structure.
By Domjan Baric, Davor Horvatic
arXiv:2606. 20560v1 Announce Type: cross Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors.
By Joshua Engels, Callum McDougall, Bilal Chughtai, Janos Kramar, Senthoran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue, Jo\~ao Gabriel Lopes de Oliveira, Rohin Shah, Neel Nanda
arXiv:2608. 01587v1 Announce Type: cross Abstract: Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows.
By Xizhe Zhang
The paper introduces SAGE, a multi‑agent framework that uses specialized analyzers to diagnose univariate time‑series anomalies by examining point, structural, seasonal, and pattern deviations. Each analyzer produces numerical evidence and visual diagnostics, which a Detector consolidates into intervals, candidate types, and confidence scores, and a Supervisor converts these into analyst‑friendly reports. Experiments on Yahoo S5, KPI, and WSD datasets show SAGE achieving the highest average Point‑F1 score (66.26) and receiving higher usefulness ratings in a blind human study.
By Hyeongwon Kang, Jeongseob Kim, Jinwoo Park, Pilsung Kang