arXiv AI

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

The paper introduces a dependency‑graph framework to formalize compositional reasoning in language models, defining three increasing levels of compositionality. Using data‑structure tasks with deterministic rewards, the authors observe a consistent asymmetry: training on decomposed skills does not reliably transfer to composed tasks, whereas training on composed tasks transfers more readily to decomposed ones. They provide a theoretical explanation for this asymmetry and evaluate its effects under length extrapolation, structural distribution shift, and transfer to unseen skills, concluding with a pilot study on real‑world tool‑calling benchmarks that suggests the phenomenon extends to practical settings.

arXiv Machine Learning
Jun 4

Can Large Language Models Generalize Procedures Across Representations?

arXiv:2602. 03542v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are often specified in natural language.

By Fangru Lin, Valentin Hofmann, Xingchen Wan, Weixing Wang, Zifeng Ding, Anthony G. Cohn, Janet B. Pierrehumbert
arXiv AI
Sep 1

Learning Composable Chains-of-Thought

arXiv:2505.22635v2 Announce Type: replace-cross Abstract: A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoni...

By Fangcong Yin, Zeyu Leo Liu, Liu Leqi, Xi Ye, Greg Durrett
arXiv Machine Learning
Jun 17

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning

arXiv:2606. 18089v1 Announce Type: new Abstract: Post-training pipelines that combine supervised fine-tuning (SFT) with reinforcement learning (RL) have emerged as the key recipe for transforming large language models (LLMs) into robust reasoners.

By Lingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma, Xiangchen Song, Yuekai Sun, Mikhail Yurochkin, Taylor W. Killian, Ruslan Salakhutdinov, Kun Zhang, Eric P. Xing, Zhengzhong Liu
arXiv AI
Sep 1

Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training

The paper investigates how different forms of compressed chain‑of‑thought (CoT) reasoning—Explicit, Composed, and Implicit—affect large language model (LLM) performance after supervised fine‑tuning (SFT). Using a synthetic compositional reasoning task, the authors show that coarser CoT requires more SFT data, that Composed and Implicit CoT benefit more from data scaling (with Composed also benefiting from repetition), and that reinforcement learning with verifiable rewards (RLVR) can decompose compressed steps learned during SFT. Additionally, unidirectional CoT ordering improves generalization on longer sequential tasks.

By Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
arXiv AI
Sep 10

Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks

The paper investigates reinforcement learning with verifiable rewards (RLVR) by expanding the reasoning space beyond depth to include environment complexity and diverse reasoning forms. It introduces a synthetic knowledge‑graph environment that varies depth, complexity, and task family, revealing that joint depth‑complexity coverage outperforms single‑axis approaches, that different reasoning families behave non‑uniformly, and that uniform mixing beats staged curricula under a fixed budget. The study also shows that current off‑the‑shelf models share a deductive‑over‑abductive bias, indicating a broader gap in reasoning capabilities.

By Yihua Zhu, Qianying Liu, Fei Cheng, Jiaxin Wang, Akiko Aizawa, Sadao Kurohashi, Hidetoshi Shimodaira
arXiv AI
Sep 15

TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

arXiv:2609.13457v1 Announce Type: new Abstract: Timeseries multimodal large language models (TS-MLLMs) have recently begun leveraging the reasoning capabilities of large language models (LLMs) for qu...

By Sudarshan Regmi, Arvind Pillai, Yu Yvonne Wu, Yuliang Chen, Bibek Panthi, Tess Z. Griffin, Michael V. Heinz, Lisa Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv AI
Aug 5

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

arXiv:2608. 03972v1 Announce Type: new Abstract: On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models.

By Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Hugging Face Trending Papers
Aug 4

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples.