arXiv Machine Learning By Tzu-Hsiang Lin, Srinivas Shakkottai, Dileep Kalathil, P. R. Kumar

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies

Read the original on arXiv Machine Learning →

arXiv:2606. 28939v1 Announce Type: new Abstract: Behavior-cloned diffusion policies are expressive but remain vulnerable to covariate shift: small deviations from demonstrated states can compound into task failure.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv AI
Sep 16

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

ReDraft is a reference‑driven revision method for continual post‑training of large multimodal language models. It uses the model’s own incorrect outputs as references, revises them, verifies the revisions, and fine‑tunes on the accepted ones, thereby combining explicit supervision with policy proximity. On tasks such as Counting, Clock Reading, and Jigsaw, ReDraft outperforms standard supervised fine‑tuning and on‑policy methods, achieving higher target‑task gains while dramatically reducing forgetting.

By Zhihao Zhang, Mingqi Wu, Qiaole Dong, Enyu Zhou, Shuo Li, Boyang Liu, Jiazheng Zhang, Honglin Guo, Xin Guo, Shaofan Liu, Junzhe Wang, Dingwei Zhu, Zhiheng Xi, Minlong Peng, Yuan Hua, Qi Zhang, Tao Gui, Xuanjing Huang