arXiv:2605. 13511v3 Announce Type: replace-cross Abstract: While many-shot ICL achieves remarkable performance, prior studies of its scaling behavior have mainly focused on non-reasoning tasks.
By Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung
arXiv:2601. 03506v2 Announce Type: replace-cross Abstract: Recent large reasoning models (LRMs) have achieved strong performance on complex reasoning tasks by generating a long chain-of-thought (Long-CoT).
By Zhaofeng Zhong, Wei Yuan, Tong Chen, Liang Qu, Xiangyu Zhao, Quoc Viet Hung Nguyen, Hongzhi Yin
The paper introduces Prototype-Mediated Process Supervision (PMPS), a method that uses learnable reasoning prototypes to provide structural supervision for latent chain-of-thought embeddings, addressing representation collapse and uneven information distribution. PMPS aligns latent and explicit CoT embeddings in a shared prototype space via many-to-many soft assignment and employs a Progressive Sequential Alignment module to guide training from positional priors to adaptive matching. Experiments show PMPS reduces output token length to under 50% of explicit CoT on GSM8K-Aug and improves accuracy by 2.08% over SIM-CoT, even surpassing CoT-SFT on GPT-2 and achieving the highest accuracy among latent reasoning methods on larger models and harder tasks.
By Yiqi Li, Xu Chen, Chen Ju, Jiangchao Yao, Zhaoyang Li, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Yu Wang
arXiv:2602. 23161v4 Announce Type: replace Abstract: Time series reasoning demands both the perception of complex dynamics and logical depth.
By Junkai Lu, Peng Chen, Xingjian Wu, Yang Shu, Chenjuan Guo, Christian S. Jensen, Bin Yang
SPEAR (Symbolic Process Evaluation and Alignment Reward) is a training‑free, plug‑and‑play reward method for on‑policy distillation in reinforcement learning. It converts natural‑language reasoning traces into domain‑adaptive symbolic milestones and uses the longest common subsequence to align student exploration with teacher milestones, producing a dense, order‑aware reward that enforces logical consistency without an external neural verifier. Experiments on math, science, and commonsense reasoning tasks show that SPEAR effectively bridges the reasoning gap between student and teacher models through sequence‑level distillation with efficient dense process rewards.
By Zhuochun Li, Yuelyu Ji, Yiming Zeng, Daqing He
arXiv:2604. 10788v2 Announce Type: replace-cross Abstract: Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools during reasoning.
By Qiancheng Xu, Yongqi Li, Fan Liu, Hongru Wang, Min Yang, Wenjie Li
PARTAB is a framework that improves large language model reasoning on tables by constructing a structured evidence interface. It represents query‑relevant evidence as semantically coherent, row‑linked table regions and performs hierarchical selection over column groups and row‑level partitions before composing the evidence for answer generation. Evaluations on multiple table reasoning benchmarks show that PARTAB consistently outperforms full‑table prompting and recent methods, achieving strong performance on WikiTableQuestions and TabFact while remaining competitive on numerical reasoning tasks.
By Md Mahadi Hasan Nahid, Davood Rafiei
arXiv:2608. 03204v1 Announce Type: cross Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks.
By Tianbao Jiang, Weicong Ni, Gerard de Melo, Linlin Wang
arXiv:2505.16782v3 Announce Type: replace
Abstract: Large Language Models (LLMs) have shown impressive performance on complex tasks through Chain-of-Thought (CoT) reasoning. However, conventional CoT...
By Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, Xiaoyu Shen
K2-V2 is a fully open, 360‑open large language model built from scratch, designed to serve as a superior base for reasoning adaptation while also supporting conversation and knowledge retrieval. It competes with leading open‑weight models in its size class, outperforming Qwen2.5‑72B and approaching Qwen3‑235B, and incorporates domain knowledge, reasoning, long‑context handling, and tool use throughout training. The authors release the complete training history, data composition, model weights, and LLM360 artifacts to enable community use and continuous training.
By K2 Team, Zhengzhong Liu, Liping Tang, Linghao Jin, Haonan Li, Nikhil Ranjan, Desai Fan, Shaurya Rohatgi, Richard Fan, Omkar Pangarkar, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Seungwook Han, Bowen Tan, Gurpreet Gosal, Xudong Han, Varad Pimpalkhute, Shibo Hao, Ming Shan Hee, Joel Hestness, Haolong Jia, Liqun Ma, Aaryamonvikram Singh, Daria Soboleva, Natalia Vassilieva, Renxi Wang, Yingquan Wu, Yuekai Sun, Taylor Killian, Alexander Moreno, John Maggs, Hector Ren, Guowei He, Hongyi Wang, Xuezhe Ma, Yuqi Wang, Mikhail Yurochkin, Eric P. Xing
MMEmb-R1 is a multimodal embedding framework that enhances reasoning by treating it as a latent variable and selecting beneficial reasoning paths through pair-aware selection and counterfactual intervention. It uses reinforcement learning to invoke reasoning only when necessary, reducing unnecessary computation and latency. On the MMEB-V2 benchmark, MMEmb-R1 achieves a state‑of‑the‑art score of 71.2 with just 4 B parameters.
By Yuchi Wang, Dingkang Yang, Haiyang Yu, Weikang Bian, Jiefeng Long, Xiao Liang, Chao Feng, Hongsheng Li
arXiv:2607. 01585v1 Announce Type: cross Abstract: Predicate invention (PI), the creation of new predicates to extend the hypothesis space, remains a critical bottleneck in Inductive Logic Programming (ILP).
By Tingting Yu, Pei-Cing Huang, Chan Hsu, Chan-Tung Ku, Yihuang Kang