arXiv AI By Yarin Bar, Yaniv Romano

Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification

Read the original on arXiv AI →

arXiv:2608. 19748v1 Announce Type: cross Abstract: Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top-ranked completion according to a reward model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 11

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.

By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
arXiv AI
Sep 3

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

The paper introduces a two‑stage training framework for compact instruction‑following rerankers. Stage 1 strengthens a 4B teacher reranker with off‑policy GRPO using LLM‑judge feedback on 88K examples, while Stage 2 trains a 1B student by sampling its own rankings and receiving soft teacher‑derived rewards, blending exploration with knowledge transfer. The method achieves superior nDCG and MRR scores on MAIR‑11 and MAIR‑Full benchmarks, outperforming offline distillation baselines and larger RL‑trained rerankers.

By Vignesh Prabhakar, Jialing Pan, Anil Babu Ankisettipalli
arXiv Machine Learning
Jul 31

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

arXiv:2607. 27787v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability.

By Ken Ding