The paper introduces LLM-Microscope, a toolkit for measuring how Large Language Models encode contextual information at the token level. It shows that seemingly minor tokens—such as determiners, stopwords, and punctuation—carry surprisingly high contextual weight, and removing them degrades performance on benchmarks like MMLU and BABILong-4k. The study also finds a strong link between contextualization and linearity, indicating that the transformation between layers can be approximated by a single linear mapping when tokens are well contextualized.
By Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov
arXiv:2606. 11371v1 Announce Type: cross Abstract: Spoken language, whether produced by humans or large language models (LLM), unfolds over time with varying semantic content.
By Han-Jen Chang, Yasir \c{C}atal, Angelika Wolman, Agust\'in Ib\'a\~nez, David Smith, I-Wen Su, Kai-Yuan Cheng, Georg Northoff
The paper introduces Coupled Usage–Sense Processes (CUSP), a method that models lexical semantic change by coupling contextual distributions through latent usage components and using Markov composition to link adjacent time periods. CUSP quantifies change magnitude and timing, separates variation into component movement and internal reorganization, and attributes changes to specific transported component pairs. The approach is validated on synthetic data, English and German corpora, and a large corpus of US court opinions, providing detailed, text‑grounded insights into how word meanings evolve over time.
By Haruka Ezoe, Ryohei Hisano
The paper studies how transformer representations evolve across layers by examining the intrinsic dimensionality (ID) of token embeddings and their neighborhood structures. It finds that closed‑class tokens expand and collapse earlier than open‑class tokens, and that these changes are linked to shifts in local geometry. The authors compare encoder and decoder models, showing distinct layer‑wise behaviors, and demonstrate that geometric features alone can predict a token’s part‑of‑speech and reveal how semantic content changes across layers.
By Samuele Vallisa, Federico Ravenda, Claudio Palominos, Rui He, Andrea Raballo, Antonietta Mira, Philipp Homan, Wolfram Hinzen
arXiv:2609.34187v2 Announce Type: replace-cross
Abstract: The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they c...
By Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer, Elisabeth Kollrack
arXiv:2608.30315v1 Announce Type: new
Abstract: Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern lang...
By Junjie Yao, Liangkai Hang, Zhi-Qin John Xu
arXiv:2606. 15521v1 Announce Type: cross Abstract: Tokenization introduces representational redundancy: under a fixed token vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string.
By Kanishk Jain, Matthew Day, Tankut Can
arXiv:2607. 12279v1 Announce Type: cross Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table.
By Jacob Dunefsky, Wes Gurnee, Emmanuel Ameisen
The paper introduces TokenAdapt, a model‑agnostic tokenizer transplantation method that uses a hybrid heuristic to initialize new token embeddings, and a novel pre‑tokenization learning approach for multi‑word Supertokens to improve compression. TokenAdapt combines local subword decomposition and global semantic similarity to preserve semantics while reducing retraining needs. Empirical results show that TokenAdapt outperforms existing baselines such as Transtokenizer and ReTok, achieving lower perplexity ratios and significant compression gains.
By Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, Adarsh Shirawalmath
arXiv:2510.22752v2 Announce Type: replace-cross
Abstract: In-context learning is governed by both temporal and semantic relationships, shaping how Large Language Models (LLMs) retrieve contextual inf...
By Anooshka Bajaj, Deven Mahesh Mistry, Sahaj Singh Maini, Yash Aggarwal, Zoran Tiganj
The paper critiques a recent NLI benchmark that tests the imperfective paradox, arguing that the benchmark suffers from conceptual and evaluation mis-specifications, notably Aspectual Reduction and a lack of strict NLI standards. The authors re-evaluate the benchmark, identify mis-specifications, and construct lexically matched minimal pairs to control for lexical variation. Their experiments reveal that models often exhibit a Sufficiency Bias, accept simple‑past hypotheses without affirming culmination, and that prompting interventions shift label decisions without improving true semantic understanding, highlighting additional failure modes such as compositional aspectual classification errors and surface‑form attraction.
By Kaiqiao Han, Yizhou Sun
MGAL is a new multilingual benchmark for evaluating long‑context large language models, built from United Nations reports in six official UN languages and covering 8K to 128K tokens. It tests four linguistic granularities—word, sentence, paragraph, and document—while also stratifying examples by their position within the document (begin, middle, end). Experiments show that models excel at word‑level tasks but struggle with coarser granularity, and that closed‑source models outperform others in lower‑resource languages, revealing challenges such as local semantic crowding and a fluency‑consistency gap.
By Chunhan Li, Chenglin Xu, Zongyang Zhang, Jiale Liu, Zhuoxi Rao, Xudong Jia, Junxiu He, Menglin Yang, Wenjuan Gong, Zhengzhe Liu, Chengwei Qin