Hugging Face Trending Papers

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

Read the original on Hugging Face Trending Papers →

Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference cost from multi-step diffusion.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
5d ago

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

UniEvo‑VL is a self‑evolving framework that lets multimodal models improve themselves by using their own critiques as privileged information. The method trains a single model to act as both teacher and student, minimizing divergence between their diffusion distributions over sampling trajectories. Experiments on Qwen‑image‑2512 show significant gains in image generation metrics, and stronger external critics further raise the improvement ceiling, though results vary across tasks.

By Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi