UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
UniEvo‑VL is a self‑evolving framework that lets multimodal models improve themselves by using their own critiques as privileged information. The method trains a single model to act as both teacher and student, minimizing divergence between their diffusion distributions over sampling trajectories. Experiments on Qwen‑image‑2512 show significant gains in image generation metrics, and stronger external critics further raise the improvement ceiling, though results vary across tasks.
Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion.
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.
arXiv:2606.18974v3 Announce Type: replace Abstract: Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly a...
arXiv:2610.02117v1 Announce Type: cross Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a froze...
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories w...