Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens. These tokens are derived from a codebook that maps embeddings to quantized visual patterns.
arXiv:2609.39688v1 Announce Type: cross
Abstract: Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment...
By Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara
Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by separate models. We study Unified Visual Generation (UVG), where a single model produces diverse image-valued outputs through a unified multimodal interface.
arXiv:2610.00341v1 Announce Type: cross
Abstract: As Large Multimodal Models (LMMs) transition toward natively unified architectures, evaluating their safety in synergistic harmful image-text generat...
By Bingjun Luo, Jialin Guo, Tony Wang, Siqi Li
InGuard introduces an inner guardrail for text-to-image generation that operates within the model’s own representations, avoiding external classifiers. It grades prompts using the text encoder’s embeddings, modifies risky embeddings with SAGE to produce safe images, and employs a latent detector to halt generation early. Evaluated on the RevGen Safety Benchmark, InGuard achieves a 97.9–98.8% safety rate across five open-weight models while reducing benign disturbances, model parameters, and denoising steps.
By Zeyu Wang, Xiaodan Li, Zhiwen Li, Yuefeng Chen, Hui Xue
The paper introduces Safety-aware Contrastive Decoding (SafeCoDe), a lightweight, model‑agnostic framework designed to improve context‑aware safety in Multimodal Large Language Models (MLLMs). SafeCoDe operates in two stages: a contrastive decoding step that highlights tokens sensitive to visual context by contrasting real and Gaussian‑noised images, and a global‑aware token modulation strategy that adjusts refusals based on scene‑level reasoning and predicted safety verdicts. Experiments across various MLLM architectures and safety benchmarks demonstrate that SafeCoDe consistently enhances context‑sensitive refusal behaviors while maintaining model helpfulness.
By Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang