ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training
Read the original on arXiv AI →ReDraft is a reference‑driven revision method for continual post‑training of large multimodal language models. It uses the model’s own incorrect outputs as references, revises them, verifies the revisions, and fine‑tunes on the accepted ones, thereby combining explicit supervision with policy proximity. On tasks such as Counting, Clock Reading, and Jigsaw, ReDraft outperforms standard supervised fine‑tuning and on‑policy methods, achieving higher target‑task gains while dramatically reducing forgetting.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.