arXiv AI
Sep 16

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

ReDraft is a reference‑driven revision method for continual post‑training of large multimodal language models. It uses the model’s own incorrect outputs as references, revises them, verifies the revisions, and fine‑tunes on the accepted ones, thereby combining explicit supervision with policy proximity. On tasks such as Counting, Clock Reading, and Jigsaw, ReDraft outperforms standard supervised fine‑tuning and on‑policy methods, achieving higher target‑task gains while dramatically reducing forgetting.

By Zhihao Zhang, Mingqi Wu, Qiaole Dong, Enyu Zhou, Shuo Li, Boyang Liu, Jiazheng Zhang, Honglin Guo, Xin Guo, Shaofan Liu, Junzhe Wang, Dingwei Zhu, Zhiheng Xi, Minlong Peng, Yuan Hua, Qi Zhang, Tao Gui, Xuanjing Huang