arXiv Computation and Language
Aug 21

SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit

arXiv:2608. 19472v1 Announce Type: new Abstract: Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing.

By Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale, Dirk Speelman
arXiv Computation and Language
Sep 22

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

The BD-LSC dataset introduces a bi‑directional lexical semantic change benchmark that tracks sense gain, loss, and stability across three time periods, while the ST‑WSD dataset offers fine‑grained, instance‑level sense annotations for words that blend slang and standard usage. These resources enable systematic evaluation of diverse models—including unsupervised clustering, supervised learning, transformer‑based approaches, and large language models—on tasks such as exact sense matching and multi‑label accuracy. The evaluation shows that few‑shot GPT‑4o performs best overall, yet all systems struggle with rare slang senses, highlighting a key open challenge in the field.

By Afnan Aloraini, Riza Batista-Navarro
arXiv Computation and Language
2d ago

Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change

The paper introduces Coupled Usage–Sense Processes (CUSP), a method that models lexical semantic change by coupling contextual distributions through latent usage components and using Markov composition to link adjacent time periods. CUSP quantifies change magnitude and timing, separates variation into component movement and internal reorganization, and attributes changes to specific transported component pairs. The approach is validated on synthetic data, English and German corpora, and a large corpus of US court opinions, providing detailed, text‑grounded insights into how word meanings evolve over time.

By Haruka Ezoe, Ryohei Hisano
arXiv Computation and Language
Sep 23

A Computational Approach to Measuring Semantic Change in Sanskrit Literature

The paper evaluates whether diachronic word embeddings can track semantic change in Sanskrit, an ancient low‑resource language with complex phonological and morphological features. A 2.7‑million‑token corpus covering four canonical periods is processed with a neural sandhi splitter and lemmatizer, and per‑period embeddings are trained. Validation against a curated set of 21 historical shifts shows that 19 shifts align with philological expectations, supporting the method’s applicability to Sanskrit.

By Tanay Agrawal
arXiv Computation and Language
Sep 17

Beyond frequency measures: Can contextual embeddings capture meaning change in scientific texts?

The study investigates whether contextual embeddings can detect meaning changes in scientific terminology beyond traditional frequency counts. Using Astrophysics and NLP corpora from 2010 to 2024, the authors extract candidate terms with KeyBERT, filter for significant frequency rises, and then evaluate semantic drift via multiple embedding‑based metrics. Results show that frequency methods slightly outperform embedding metrics in aligning with expert judgments, yet embedding‑only detections (e.g., "primordial black holes") reveal critical conceptual shifts missed by frequency alone, suggesting complementary value.

By Jianying Liu (STL, BETA, CEIPI), Kim Gerdes (LISN, Qatent, STL), Jean-Marc Deltorn (CEIPI)