arXiv Machine Learning By Lingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma, Xiangchen Song, Yuekai Sun, Mikhail Yurochkin, Taylor W. Killian, Ruslan Salakhutdinov, Kun Zhang, Eric P. Xing, Zhengzhong Liu

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning

Read the original on arXiv Machine Learning →

arXiv:2606. 18089v1 Announce Type: new Abstract: Post-training pipelines that combine supervised fine-tuning (SFT) with reinforcement learning (RL) have emerged as the key recipe for transforming large language models (LLMs) into robust reasoners.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 1

Learning Composable Chains-of-Thought

arXiv:2505.22635v2 Announce Type: replace-cross Abstract: A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoni...

By Fangcong Yin, Zeyu Leo Liu, Liu Leqi, Xi Ye, Greg Durrett
arXiv AI
Sep 1

Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training

The paper investigates how different forms of compressed chain‑of‑thought (CoT) reasoning—Explicit, Composed, and Implicit—affect large language model (LLM) performance after supervised fine‑tuning (SFT). Using a synthetic compositional reasoning task, the authors show that coarser CoT requires more SFT data, that Composed and Implicit CoT benefit more from data scaling (with Composed also benefiting from repetition), and that reinforcement learning with verifiable rewards (RLVR) can decompose compressed steps learned during SFT. Additionally, unidirectional CoT ordering improves generalization on longer sequential tasks.

By Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
arXiv AI
Sep 18

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

The paper introduces a dependency‑graph framework to formalize compositional reasoning in language models, defining three increasing levels of compositionality. Using data‑structure tasks with deterministic rewards, the authors observe a consistent asymmetry: training on decomposed skills does not reliably transfer to composed tasks, whereas training on composed tasks transfers more readily to decomposed ones. They provide a theoretical explanation for this asymmetry and evaluate its effects under length extrapolation, structural distribution shift, and transfer to unseen skills, concluding with a pilot study on real‑world tool‑calling benchmarks that suggests the phenomenon extends to practical settings.

By Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik
arXiv AI
Aug 18

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

arXiv:2604. 06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes.

By Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu