TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
arXiv:2606. 01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains.
arXiv:2608. 14270v1 Announce Type: new Abstract: Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions.
arXiv:2606. 01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains.
TimeEvo is a new method for time‑series agents that autonomously evolves its tool library based on failures observed during runtime. By clustering diagnosed failures into capability gaps, planning measurements, synthesizing evidence‑only tools, and admitting candidates through a paired gate, the system starts from an empty library and improves accuracy across ten QA tasks and three backbones. Experiments show that even a library built on a cheap model benefits stronger models when installed.
arXiv:2609.24862v1 Announce Type: new Abstract: Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning s...
ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.
Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time...
arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...
The study evaluates how large language model agents maintain consistency over extended interactions by simulating a 20‑step delayed‑gratification task. Researchers ran 84,540 trajectories across eight model families, using survival analysis to track when agents first claim a reward and discrete‑time hazard regression to assess how factors like social visibility, persona stressors, and deliberation policy affect failure risk. They also developed a seven‑category taxonomy from 13,780 deliberation traces, revealing that early failures are impulse‑driven, later ones are fatigue‑ or cost‑benefit‑framed, and public settings elicit norm‑oriented justifications; longer deliberation correlates with higher intra‑rationale contradictions, challenging assumptions about reasoning depth and consistency.
arXiv:2509. 11575v3 Announce Type: replace Abstract: Time series reasoning treats time as a first-class axis and incorporates intermediate evidence directly into the answer.
EvoTS-Agent is a self‑evolving large language model agent designed for autonomous change‑point detection in financial time series. It begins with curated exploratory data analysis to set up candidate models, then iteratively refines detection pipelines using three operators—Revision, Alternative Strategy, and Recombination—guided by validation feedback. Across four benchmark datasets, EvoTS-Agent consistently outperforms existing LLM‑based agents and achieves a 100% execution success rate with all tested backbone LLMs.
The paper investigates when forecasting agents should employ different behaviors—retrieval, reasoning, deferring to market priors, or using historical analogs—on binary forecasting tasks. It finds that the optimal mechanism depends on the data source, with structured analogs excelling for some processes and market or conservative baselines for others. The authors propose ReliabilityRoute, a rule‑based system that steers agent behavior using reliability features, achieving competitive performance across multiple LLM versions while highlighting that more reasoning is not always better.
arXiv:2606. 05404v1 Announce Type: cross Abstract: Time series are often embedded in rich contexts that are essential for holistic modeling.
LongDS-Bench is a new benchmark for evaluating long-horizon, multi-turn data analysis by agents, featuring 68 tasks derived from real-world Kaggle notebooks that span 2,225 turns across six domains such as Geoscience, Business, and Education. The benchmark focuses on agents’ ability to maintain, update, restore, and compose evolving analytical states, with tasks designed around state-evolution patterns like counterfactual perturbation, rollback, and multi-state composition, and an average dependency span of 11.3 turns. Evaluation of five state-of-the-art models shows that the best model achieves only 48.45% average accuracy, with performance dropping nearly 47 points from early to late turns and long-horizon errors accounting for 52%–69% of failures, indicating that maintaining a correct analytical state is the key bottleneck.