arXiv:2602. 00462v4 Announce Type: replace-cross Abstract: Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM.
By Benno Krojer, Shravan Nayak, Oscar Ma\~nas, Vaibhav Adlakha, Desmond Elliott, Siva Reddy, Marius Mosbach
The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.
By Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov
arXiv:2608.23551v1 Announce Type: cross
Abstract: Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existi...
By Na Li, Yuchen Jiao, Changxiao Cai, Gen Li
arXiv:2605. 10938v2 Announce Type: replace-cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.
By Keya Hu, Linlu Qiu, Yiyang Lu, Hanhong Zhao, Tianhong Li, Yoon Kim, Jacob Andreas, Kaiming He
arXiv:2606.03715v3 Announce Type: replace
Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that co...
By Nurit Spingarn, Noa Cohen, Tamar Rott Shaham, Tomer Michaeli
Diffusion Trajectory Modeling (DTM) treats the evolving feature maps of diffusion models as temporally structured trajectories rather than static snapshots. By interpreting each spatial patch’s progression across multiple timesteps as a trajectory, DTM captures semantic correspondence cues that prior methods miss. Experiments on SPair-71k, SPair-U, and AP-10K demonstrate that DTM achieves strong performance, highlighting the semantic value embedded in the diffusion process’s temporal axis.
By Yusung Choi
ORCA (Orthogonal Residual Compositional Alignment) is a method that improves text-to-image diffusion models by aligning the diffusion transformer’s latent space with a low‑rank target derived from a frozen visual encoder. It introduces an auxiliary loss that uses a predictor with an orthogonal basis parameterised by a learned residual between T5 and CLIP embeddings, providing a prompt‑dependent signal for selecting the visual readout subspace. Experiments on three diffusion‑transformer backbones show that ORCA improves FID and GenEval scores, especially on attribute binding, spatial relations, and multi‑object prompts, without adding inference‑time cost.
By Arshia Hemmat, Amirhossein Vahidi, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, Mohammad Lotfollahi
arXiv:2602.10216v2 Announce Type: replace
Abstract: A single text prompt passed to a diffusion model yields a wide range of visual outputs determined solely by a stochastic process, leaving users wit...
By Pawe{\l} Skier\'s, Emilia Kaczmarczyk, Tomasz Trzci\'nski, Kamil Deja
arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.
By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia
arXiv:2606. 08347v1 Announce Type: cross Abstract: Modern language models represent text using discrete token-level embeddings, which forces recurring multi-token patterns to be learned implicitly across Transformer layers.
By Wuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Yuning Qiu, Qibin Zhao, Danilo Mandic
arXiv:2609.37225v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods eith...
By Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong
CausalEmbed is an auto‑regressive method for generating compact multi‑vector embeddings in visual document retrieval. By using iterative margin loss during contrastive training, it reduces the number of visual tokens needed by 30‑155× while keeping performance competitive across different backbones and benchmarks. The approach offers efficient training, scalable test‑time performance, and a flexible scaling strategy for multi‑vector representations.
By Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu