arXiv Machine Learning By Thomas Fabian

Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging

Read the original on arXiv Machine Learning →

arXiv:2601. 05713v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) represent natural language is a central challenge in natural language processing (NLP) research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 15

LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.

By Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov
arXiv AI
Jun 29

ELF: Embedded Language Flows

arXiv:2605. 10938v2 Announce Type: replace-cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.

By Keya Hu, Linlu Qiu, Yiyang Lu, Hanhong Zhao, Tianhong Li, Yoon Kim, Jacob Andreas, Kaiming He
arXiv Computer Vision
Sep 15

Diffusion Trajectory Modeling for Semantic Correspondence

Diffusion Trajectory Modeling (DTM) treats the evolving feature maps of diffusion models as temporally structured trajectories rather than static snapshots. By interpreting each spatial patch’s progression across multiple timesteps as a trajectory, DTM captures semantic correspondence cues that prior methods miss. Experiments on SPair-71k, SPair-U, and AP-10K demonstrate that DTM achieves strong performance, highlighting the semantic value embedded in the diffusion process’s temporal axis.

By Yusung Choi