arXiv Machine Learning

Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging

arXiv:2601. 05713v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) represent natural language is a central challenge in natural language processing (NLP) research.

arXiv AI
Sep 15

LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.

By Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov
arXiv AI
Jun 29

ELF: Embedded Language Flows

arXiv:2605. 10938v2 Announce Type: replace-cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.

By Keya Hu, Linlu Qiu, Yiyang Lu, Hanhong Zhao, Tianhong Li, Yoon Kim, Jacob Andreas, Kaiming He
arXiv Computer Vision
Sep 15

Diffusion Trajectory Modeling for Semantic Correspondence

Diffusion Trajectory Modeling (DTM) treats the evolving feature maps of diffusion models as temporally structured trajectories rather than static snapshots. By interpreting each spatial patch’s progression across multiple timesteps as a trajectory, DTM captures semantic correspondence cues that prior methods miss. Experiments on SPair-71k, SPair-U, and AP-10K demonstrate that DTM achieves strong performance, highlighting the semantic value embedded in the diffusion process’s temporal axis.

By Yusung Choi
arXiv Machine Learning
1d ago

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

ORCA (Orthogonal Residual Compositional Alignment) is a method that improves text-to-image diffusion models by aligning the diffusion transformer’s latent space with a low‑rank target derived from a frozen visual encoder. It introduces an auxiliary loss that uses a predictor with an orthogonal basis parameterised by a learned residual between T5 and CLIP embeddings, providing a prompt‑dependent signal for selecting the visual readout subspace. Experiments on three diffusion‑transformer backbones show that ORCA improves FID and GenEval scores, especially on attribute binding, spatial relations, and multi‑object prompts, without adding inference‑time cost.

By Arshia Hemmat, Amirhossein Vahidi, Amitis Shidani, Mohammad Vali Sanian, Hesam Asadollahzadeh, Aryan Yazdan Parast, Mohammad Lotfollahi
arXiv AI
Jul 8

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.

By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia
arXiv Computation and Language
Sep 23

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

CausalEmbed is an auto‑regressive method for generating compact multi‑vector embeddings in visual document retrieval. By using iterative margin loss during contrastive training, it reduces the number of visual tokens needed by 30‑155× while keeping performance competitive across different backbones and benchmarks. The approach offers efficient training, scalable test‑time performance, and a flexible scaling strategy for multi‑vector representations.

By Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu