arXiv:2606. 01591v1 Announce Type: cross Abstract: The TimeLogic Challenge evaluates formal temporal-logic reasoning over video - 16 operators (before, after, until, since, always, co-occur, ordering, ...
By Ali Alavi
arXiv:2607.22365v2 Announce Type: replace
Abstract: High-complexity operational environments require methods that characterize temporally distributed patterns rather than classify isolated events. Th...
By Michael Romei De Socio, Gian Luca Pozzato, Alessio Merlo
The paper proposes PAIR, a method that treats reasoning paths of large language models as phase‑structured trajectories within each question. By sampling multiple trajectories per question, aligning them to shared relative phases, and comparing successful versus unsuccessful paths only within the same phase, PAIR isolates path‑quality signals from question‑level variation. Experiments show that standard correctness probes lose predictive power under this within‑question evaluation, while PAIR improves trajectory ranking, Best‑of‑N selection, and enables phase‑wise steering of generation outcomes.
By Zhenghao He, Guangzhi Xiong, Sanchit Sinha, Bohan Liu, Wenqian Ye, Aidong Zhang
arXiv:2609.16055v1 Announce Type: cross
Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasonin...
By Zhiren Gong, Yikun Hou, Zihao Zeng, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim
arXiv:2609.22213v1 Announce Type: new
Abstract: Temporal Knowledge Graph Question Answering (TKGQA) requires answer inference from evidence that is both structurally valid and temporally admissible....
By Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Dongjin Yu, Yu Wang
arXiv:2606. 12481v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong reasoning and instruction-following capabilities, making them potentially powerful tools for time-series analysis.
By Jaeho Kim, Changhun Oh, Seokhyun Lee, Irina Rish, Changhee Lee
arXiv:2609.13457v1 Announce Type: new
Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for qu...
By Sudarshan Regmi, Arvind Pillai, Yu Yvonne Wu, Yuliang Chen, Bibek Panthi, Tess Z. Griffin, Michael V. Heinz, Lisa Marsch, Nicholas C. Jacobson, Andrew Campbell
A*-Thought-V2 is a framework that models Chain-of-Thought reasoning as a geometric trajectory in a 3D PCA space, using explicit-implicit latent tokens to compress steps that deviate from the main question-to-solution direction. The method measures alignment angles to decide which steps remain text and which become latent, and introduces stepwise embedding forcing and label forcing to train the architecture. Experiments on Qwen models show up to 2.6% accuracy gains, halved response length, and significant reductions in computation and training time.
By Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
arXiv:2607. 20560v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems retrieve and integrate external knowledge to ground large language model (LLM) outputs.
By Muntaser Syed, Marius Silaghi, Sheikh Abujar, Sharun Akter
arXiv:2608.20359v1 Announce Type: new
Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performanc...
By Ravisri Valluri, Tung Nguyen, Aditya Grover
arXiv:2609.14142v1 Announce Type: new
Abstract: Large language models (LLMs) can struggle with time-series question answering (TS-QA), especially when numerical signals are serialized as text and req...
By Ivan Delgado, Himansi Gupta, Bishal Khatri, Niharika Sapre, Lameta Shamoon, Onat Gungor, Tajana Rosing
arXiv:2606. 05402v1 Announce Type: cross Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process.
By Jinu Lee, Shivam Agarwal, Amruta Parulekar, Siddarth Madala, Dilek Hakkani-Tur, Julia Hockenmaier