arXiv AI By Shunyu Wu, Dan Li, Haozheng Ye, Weibin Feng, Jian Lou, Bo Zhang, Wenjie Feng, Chenjuan Guo, See-Kiong Ng

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

Read the original on arXiv AI →

arXiv:2606. 03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

TSQueryBench is a synthetic benchmark comprising 500 time‑series instances and 10 query types, each paired with correct, partially correct, and incorrect natural‑language explanations. The study evaluates six large language models on explanation generation, ranking, scoring, and anomaly detection, revealing that models often fail to generate numerically correct explanations yet can reliably identify or score correct ones. These findings suggest that rubric‑guided LLM evaluation is more dependable than generation for numerically grounded time‑series reasoning.

By Preetham Sivalingam, Murari Mandal, Dhruv Kumar, Saurabh Deshpande
arXiv AI
Jun 2

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.

By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller
arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi