arXiv Computation and Language
Sep 18

WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution

The paper investigates why Word-in-Context (WiC) remains difficult for language models, suggesting that the lack of an explicit sense inventory contributes to the challenge. By evaluating open LLMs on both WiC and traditional Word Sense Disambiguation (WSD) tasks, the authors find that providing candidate senses—akin to WSD—consistently improves WiC performance. Human evaluation indicates that many WiC errors stem from label ambiguity or mismatched sense boundaries, with models often over‑discriminating senses and making overly fine‑grained distinctions.

By Yi Zhou, Kiamehr Rezaee, Danushka Bollegala, Mohammad Taher Pilehvar, Jose Camacho-Collados
arXiv AI
6d ago

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

The paper introduces TokenAdapt, a model‑agnostic tokenizer transplantation method that uses a hybrid heuristic to initialize new token embeddings, and a novel pre‑tokenization learning approach for multi‑word Supertokens to improve compression. TokenAdapt combines local subword decomposition and global semantic similarity to preserve semantics while reducing retraining needs. Empirical results show that TokenAdapt outperforms existing baselines such as Transtokenizer and ReTok, achieving lower perplexity ratios and significant compression gains.

By Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, Adarsh Shirawalmath
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani