TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
arXiv:2606. 01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains.
arXiv:2606. 15107v1 Announce Type: new Abstract: Time series data in real-world deployments is overwhelmingly irregular.
arXiv:2606. 01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains.
arXiv:2606. 05404v1 Announce Type: cross Abstract: Time series are often embedded in rich contexts that are essential for holistic modeling.
arXiv:2609.13457v1 Announce Type: new Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for qu...
arXiv:2606. 02433v1 Announce Type: cross Abstract: The rapid development of LLMs has significantly advanced tabular question answering, but most systems cannot perform future-oriented numerical prediction.
TSQueryBench is a synthetic benchmark comprising 500 time‑series instances and 10 query types, each paired with correct, partially correct, and incorrect natural‑language explanations. The study evaluates six large language models on explanation generation, ranking, scoring, and anomaly detection, revealing that models often fail to generate numerically correct explanations yet can reliably identify or score correct ones. These findings suggest that rubric‑guided LLM evaluation is more dependable than generation for numerically grounded time‑series reasoning.
arXiv:2606. 10460v1 Announce Type: cross Abstract: Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved.
arXiv:2509. 11575v3 Announce Type: replace Abstract: Time series reasoning treats time as a first-class axis and incorporates intermediate evidence directly into the answer.
arXiv:2606. 12481v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong reasoning and instruction-following capabilities, making them potentially powerful tools for time-series analysis.
TimeEvo is a new method for time‑series agents that autonomously evolves its tool library based on failures observed during runtime. By clustering diagnosed failures into capability gaps, planning measurements, synthesizing evidence‑only tools, and admitting candidates through a paired gate, the system starts from an empty library and improves accuracy across ten QA tasks and three backbones. Experiments show that even a library built on a cheap model benefits stronger models when installed.
arXiv:2608. 03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features.
arXiv:2603. 02668v2 Announce Type: replace Abstract: We present SorryDB, a dynamically-updating benchmark of open Lean tasks drawn from 78 real world formalization projects on GitHub.
arXiv:2607. 06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored.