arXiv AI By Zixuan Wang, Xingyu Dang, Jason D. Lee, Kaifeng Lyu

The Power of Power Law: Asymmetry Enables Compositional Reasoning

Read the original on arXiv AI →

arXiv:2604. 22951v2 Announce Type: replace Abstract: Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

Learning Composable Chains-of-Thought

arXiv:2505.22635v2 Announce Type: replace-cross Abstract: A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoni...

By Fangcong Yin, Zeyu Leo Liu, Liu Leqi, Xi Ye, Greg Durrett
arXiv AI
Sep 18

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

The paper introduces a dependency‑graph framework to formalize compositional reasoning in language models, defining three increasing levels of compositionality. Using data‑structure tasks with deterministic rewards, the authors observe a consistent asymmetry: training on decomposed skills does not reliably transfer to composed tasks, whereas training on composed tasks transfers more readily to decomposed ones. They provide a theoretical explanation for this asymmetry and evaluate its effects under length extrapolation, structural distribution shift, and transfer to unseen skills, concluding with a pilot study on real‑world tool‑calling benchmarks that suggests the phenomenon extends to practical settings.

By Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik
arXiv AI
Sep 7

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

The paper investigates data efficiency and selection in On‑Policy Distillation (OPD) for large language models. It shows that 1‑shot OPD—training on a single example—consistently improves performance, especially when the example is hard, and that longer chain‑of‑thought (CoT) paths drive the gains rather than token entropy. Based on these findings, the authors propose a simple hard‑example selection strategy that, using only eight carefully chosen hard examples, matches the performance of a 17,000‑example baseline across models from 1.5B to 7B parameters.

By Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You