arXiv Computation and Language
2d ago

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

The paper proposes a new method for training large language models to handle long-context reasoning by combining Group Relative Policy Optimization (GRPO) with on‑policy distillation (OPD). It introduces a synthetic multilingual dataset called LongBlocks that tests multi‑hop reasoning, contextual grounding, and long‑form generation. Experiments show that the combined approach outperforms either GRPO or OPD alone while maintaining short‑context performance.

By Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins