arXiv AI By Lin Li, Jiawei Huang, Qihao Quan, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Wenjie Feng, Jian Lou, See-Kiong Ng

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 11

KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning

KairosAgent is an agentic framework that combines a large language model (LLM) reasoner with a time series foundation model (TSFM) forecaster to tackle cross‑domain multimodal time series forecasting. It dynamically invokes analytical tools to improve the LLM’s numerical comprehension and semantic reasoning, then fuses the reasoning outcomes into the TSFM pipeline for more accurate predictions. The approach is further enhanced by a curated large‑scale trajectory corpus and a reinforcement learning paradigm with multi‑turn refinement and turn‑level credit assignment, achieving superior zero‑shot forecasting performance.

By Kun Feng, Ziwei Shan, Yuchen Fang, Yiyang Tan, Sihan Lu, Shuqi Gu, Xingyu Lu, Lintao Ma, Kan Ren
arXiv Machine Learning
Sep 24

ChronoSteer: Bridging Large Language Model and Time Series Foundation Model via Synthetic Cross-Modal Alignment Dataset

ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.

By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
arXiv Computer Vision
Aug 26

DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton

The paper introduces DoublesEval, a diagnostic framework that uses professional doubles badminton to test visual‑language models’ ability to reason about dynamic multi‑agent interactions. It decomposes rallies into key moments and evaluates models across four dimensions—atomic recognition, intra‑segment composite understanding, cross‑segment causal reasoning, and high‑level tactical abstraction—highlighting specific reasoning failures. The authors also propose TacticCheck, a lightweight consistency checker that improves performance without retraining the models, yet significant gaps remain in tactical reasoning.

By Jintao Cheng, Weibin Li