arXiv Computer Vision

ProgressNet: Sketching and Prompting with a Frozen Text-to-Image Model

ProgressNet is a training‑free framework that enables a frozen text‑to‑image model to follow a sketching session in real time. It allows strokes to be added or erased, prompts to be revised, and produces an updated image in about a second per turn without adding new parameters. The method uses three inference‑time mechanisms—Previous‑Concept Memory, Layer‑Selective K/V Injection, and Banded Adaptive Control—to maintain fidelity and coherence across sketch domains, outperforming existing approaches especially when sketches are partially erased.

arXiv AI
Sep 2

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

The paper introduces GLANCE, a one‑pass block drafting method that enables lossless speculative decoding for vision‑language models. By using a block‑diffusion head that reads the fused vision‑language state, GLANCE eliminates the need for the drafter to process the image at every step, allowing it to fill an entire block in a single forward pass. Experiments show that GLANCE can decode up to 2.93× faster than autoregressive decoding while maintaining exact greedy decoding results across multiple tasks.

By Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
arXiv AI
Sep 15

Freeze, Share, Shrink: Rethinking the Action Backbone in Diffusion Policies

The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.

By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU
Hugging Face Trending Papers
Jun 11

InterleaveThinker: Reinforcing Agentic Interleaved Generation

Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation.

arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv Computer Vision
Sep 11

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

AcFlow introduces an inference‑time controller for text‑to‑image diffusion transformers that transports intermediate layer activations through a learned, concept‑conditioned velocity field while keeping the base model frozen. The method allows fine‑grained style intensity control and suppression of unwanted concepts, achieving superior style–content trade‑offs compared to baselines and generalizing to unseen concepts without per‑concept fitting. Experiments demonstrate improved style alignment and qualitative suppression of diverse concepts where direct prompting fails.

By Junran Wang, Zehao Jin, Tianyu Luan, Xinjie Shen
arXiv Computer Vision
Sep 2

Training-Free Inpainting Across Domains with a Frozen Text-to-Image Diffusion Model

The paper demonstrates that a frozen generic text‑to‑image diffusion model can perform conditional inpainting across three natural‑image domains without any inpainting‑specific training or dataset adaptation. The proposed Step‑PI method augments known‑region projection with boundary‑interior latent feedback, persistent PI state, and a predefined four‑field release schedule, improving all evaluated metrics over baseline training‑free approaches. Experiments on CelebA‑HQ, AFHQ, and Places2 show consistent gains, with Step‑PI outperforming LanPaint and PILOT on all five macro metrics.

By Zhenhuan Wang, Fengyi Yuan
arXiv Computer Vision
Sep 4

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics

The paper introduces a lightweight model that learns controlled latent dynamics to accelerate masked image generation models (MIGMs). By incorporating previous features and sampled tokens, it regresses the average velocity field of feature evolution, reducing redundancy from bi-directional attention. Applied to Lumina-DiMOO, the method achieves over 4× faster text-to-image generation while preserving quality, advancing the efficiency frontier for MIGMs.

By Kaiwen Zhu, Quansheng Zeng, Yuandong Pu, Shuo Cao, Xiaohui Li, Yi Xin, Qi Qin, Jiayang Li, Juncheng Yan, Yu Qiao, Jinjin Gu, Yihao Liu
arXiv Computer Vision
Oct 2

RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models

RASteer is a training‑free method for concept erasure in text‑to‑image diffusion models that retains other concepts. It constructs a retain subspace from concepts to preserve, then removes components aligned with this subspace from the erasure direction using Retain‑Orthogonal Steering (ROS). Overlap‑Adaptive Calibration (OAC) further adjusts the removal of shared components at each layer and denoising step, balancing target erasure with concept preservation. Experiments show RASteer matches or outperforms existing activation steering and weight editing baselines on unsafe‑content, instance, and artistic‑style erasure tasks across multiple backbones and benchmarks.

By Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu
arXiv AI
Aug 7

CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

arXiv:2603. 08652v2 Announce Type: replace Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning.

By Haodong Li, Chunmei Qing, Huanyu Zhang, Dongzhi Jiang, Yihang Zou, Hongbo Peng, Dingming Li, Yuhong Dai, ZePeng Lin, Juanxi Tian, Yi Zhou, Siqi Dai, Jingwei Wu, Pheng-Ann Heng