TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
arXiv:2606. 01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains.
arXiv:2606. 03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions.
arXiv:2606. 01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains.
TSQueryBench is a synthetic benchmark comprising 500 time‑series instances and 10 query types, each paired with correct, partially correct, and incorrect natural‑language explanations. The study evaluates six large language models on explanation generation, ranking, scoring, and anomaly detection, revealing that models often fail to generate numerically correct explanations yet can reliably identify or score correct ones. These findings suggest that rubric‑guided LLM evaluation is more dependable than generation for numerically grounded time‑series reasoning.
arXiv:2601. 17717v3 Announce Type: replace Abstract: Large Language Models (LLMs) have emerged as powerful tools for generating data across various modalities.
arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.
arXiv:2607. 28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples.
PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.
arXiv:2604. 11996v2 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy?
arXiv:2608. 06329v1 Announce Type: cross Abstract: Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed.
arXiv:2607. 01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society.
arXiv:2603. 00546v2 Announce Type: replace Abstract: Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains.
arXiv:2509. 11575v3 Announce Type: replace Abstract: Time series reasoning treats time as a first-class axis and incorporates intermediate evidence directly into the answer.
arXiv:2606. 11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems.