HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments
Read the original on arXiv AI →HISPO introduces a segment‑level policy‑optimization technique for reinforcement learning with verifiable rewards, creating entropy‑derived contiguous segments during rollout and applying clipped importance‑sampling correction at this granularity. It bridges the gap between token‑level and sequence‑level corrections, offering a middle‑ground approach for credit assignment in long‑form mathematical reasoning. Evaluated on Qwen3‑1.7B‑Base across six benchmarks, HISPO consistently outperforms or matches the strongest baselines in Pass@8 and Acc@8 metrics, notably improving AIME25 scores over GRPO and GSPO.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.