arXiv Machine Learning

Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language

arXiv:2608. 05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap.

arXiv AI
Sep 10

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.

By Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
arXiv Computation and Language
3d ago

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...

By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
arXiv AI
2d ago

LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

LineupRL introduces a reinforcement learning framework with verifiable rewards for time series captioning, using a frozen large language model to identify the correct time series from a set of distractors based on a generated caption. This approach bypasses the limitations of supervised fine‑tuning and traditional RL rewards that poorly transfer to open‑ended time series generation. Experiments on two captioning benchmarks, as well as forecasting and reconstruction tasks, show that LineupRL outperforms both SFT and RL baselines across all metrics, and its trained 3B vision‑language model surpasses a 72B model distilled from SFT captions. The method also demonstrates resistance to reward hacking and produces captions that accurately trace trends and name key values.

By Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen
Hugging Face Trending Papers
Aug 10

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.