The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.
By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.
By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation.
The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.
By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen
The paper demonstrates that a frozen generic text‑to‑image diffusion model can perform conditional inpainting across three natural‑image domains without any inpainting‑specific training or dataset adaptation. The proposed Step‑PI method augments known‑region projection with boundary‑interior latent feedback, persistent PI state, and a predefined four‑field release schedule, improving all evaluated metrics over baseline training‑free approaches. Experiments on CelebA‑HQ, AFHQ, and Places2 show consistent gains, with Step‑PI outperforming LanPaint and PILOT on all five macro metrics.
By Zhenhuan Wang, Fengyi Yuan
arXiv:2606. 31495v1 Announce Type: new Abstract: We study a single idea across two settings: that a prediction-error signal, computed by a small predictor over the latent space of a frozen encoder, can serve both as a gate on plasticity and as a substrate for metacognition.
By Louis Mouchon
The paper introduces a lightweight model that learns controlled latent dynamics to accelerate masked image generation models (MIGMs). By incorporating previous features and sampled tokens, it regresses the average velocity field of feature evolution, reducing redundancy from bi-directional attention. Applied to Lumina-DiMOO, the method achieves over 4× faster text-to-image generation while preserving quality, advancing the efficiency frontier for MIGMs.
By Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin, Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu
RASteer is a training‑free method for concept erasure in text‑to‑image diffusion models that retains other concepts. It constructs a retain subspace from concepts to preserve, then removes components aligned with this subspace from the erasure direction using Retain‑Orthogonal Steering (ROS). Overlap‑Adaptive Calibration (OAC) further adjusts the removal of shared components at each layer and denoising step, balancing target erasure with concept preservation. Experiments show RASteer matches or outperforms existing activation steering and weight editing baselines on unsafe‑content, instance, and artistic‑style erasure tasks across multiple backbones and benchmarks.
By Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu
arXiv:2603. 08652v2 Announce Type: replace Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning.
By Haodong Li, Chunmei Qing, Huanyu Zhang, Dongzhi Jiang, Yihang Zou, Hongbo Peng, Dingming Li, Yuhong Dai, ZePeng Lin, Juanxi Tian, Yi Zhou, Siqi Dai, Jingwei Wu, Pheng-Ann Heng
arXiv:2604.22875v3 Announce Type: replace-cross
Abstract: When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-languag...
By Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri, Trung Bui, Anh Totti Nguyen
arXiv:2610.00623v1 Announce Type: new
Abstract: Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where li...
By Wenhan Yang, Anirudh Rao, Ashwin Chandra