ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.
By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
BiMoGen introduces a unified masked discrete diffusion framework for bidirectional motion‑text generation, addressing the limitations of autoregressive models in capturing bidirectional dependencies between language and motion. The approach employs a two‑stage training strategy—decoupled uni‑ and cross‑modal pretraining followed by supervised fine‑tuning—to establish robust cross‑modal correspondence, and incorporates Generation‑Aware Self‑Correction to mitigate error propagation during inference. Experiments on HumanML3D and KIT‑ML show competitive performance on both text‑to‑motion and motion‑to‑text tasks, demonstrating the effectiveness of the proposed training and correction mechanisms.
NeST is a framework that adapts large language models (LLMs) for continuous time‑series forecasting by creating neighborhood‑aware text prototypes and aligning them with temporal representations through a nearest‑neighbor contrastive objective. It retrieves the most relevant prototypes and uses them to conditionally modulate time‑series features, enabling more effective integration of textual and temporal information. Experiments show that NeST outperforms state‑of‑the‑art methods on eight benchmarks, reduces MSE by 1.2% for long‑term forecasting, improves zero‑shot forecasting by 4.9%, and boosts R² by 3.3% on a real‑world photovoltaic power forecasting task.
By Jayanie Bogahawatte, Sachith Seneviratne, Maneesha Perera, Saman Halgamuge
arXiv:2509. 07295v4 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture.
By Ji Xie, Trevor Darrell, Luke Zettlemoyer, XuDong Wang
arXiv:2607. 00858v1 Announce Type: cross Abstract: Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations.
By Peiyuan Zhu, Shaoan Xie, Zijian Li, Yifan Shen, Namrata Deka, Harsh Shrivastava, Guangyi Chen, Kun Zhang
arXiv:2609.00505v1 Announce Type: new
Abstract: Video-text models adapted from image-text architectures (e.g., CLIP) frequently exhibit temporal blindness, the inability to perceive fundamental cues...
By Sethuraman T V, Savya Khosla, Onkar Kishor Susladkar, Aditi Tiwari, Seoung Wug Oh, Kushal Kafle, Joon-Young Lee, Derek Hoiem, Simon Jenni