arXiv AI By Chenhao Qiu, Dawei Li, Yechao Zhang, Lei Gong, Zhen Tan

OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

Read the original on arXiv AI →

OPD‑Aha is a privileged on‑policy distillation method that improves multimodal reasoning by reconstructing the distillation target from the teacher’s isolated visual preference instead of relying on fragile teacher‑student discrepancies. It suppresses continuations that contradict the image, encouraging students to interrupt flawed reasoning with reflection tokens such as "wait" and "actually." This approach leads to consistent improvements across fine‑grained perception and complex multimodal reasoning benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.