arXiv Computation and Language

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

The BD-LSC dataset introduces a bi‑directional lexical semantic change benchmark that tracks sense gain, loss, and stability across three time periods, while the ST‑WSD dataset offers fine‑grained, instance‑level sense annotations for words that blend slang and standard usage. These resources enable systematic evaluation of diverse models—including unsupervised clustering, supervised learning, transformer‑based approaches, and large language models—on tasks such as exact sense matching and multi‑label accuracy. The evaluation shows that few‑shot GPT‑4o performs best overall, yet all systems struggle with rare slang senses, highlighting a key open challenge in the field.

arXiv Computation and Language
Aug 21

SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit

arXiv:2608. 19472v1 Announce Type: new Abstract: Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing.

By Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale, Dirk Speelman
arXiv Computation and Language
2d ago

Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change

The paper introduces Coupled Usage–Sense Processes (CUSP), a method that models lexical semantic change by coupling contextual distributions through latent usage components and using Markov composition to link adjacent time periods. CUSP quantifies change magnitude and timing, separates variation into component movement and internal reorganization, and attributes changes to specific transported component pairs. The approach is validated on synthetic data, English and German corpora, and a large corpus of US court opinions, providing detailed, text‑grounded insights into how word meanings evolve over time.

By Haruka Ezoe, Ryohei Hisano
arXiv Computation and Language
Sep 23

From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits

The paper investigates how internet slang spreads across Reddit communities by combining social network analysis with linguistic context. Using large language models as scalable annotators, the authors create a benchmark for detecting slang usage and then model its adoption and diffusion. Findings reveal that users with higher bridging capital promote slang spread, while those with higher bonding capital hinder it, and that broader contextual usage delays new user adoption.

By Xiaoning Wang, Ted Underwood, Zhewei Sun
arXiv Computation and Language
Sep 18

WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution

The paper investigates why Word-in-Context (WiC) remains difficult for language models, suggesting that the lack of an explicit sense inventory contributes to the challenge. By evaluating open LLMs on both WiC and traditional Word Sense Disambiguation (WSD) tasks, the authors find that providing candidate senses—akin to WSD—consistently improves WiC performance. Human evaluation indicates that many WiC errors stem from label ambiguity or mismatched sense boundaries, with models often over‑discriminating senses and making overly fine‑grained distinctions.

By Yi Zhou, Kiamehr Rezaee, Danushka Bollegala, Mohammad Taher Pilehvar, Jose Camacho-Collados
arXiv Computation and Language
Sep 17

Beyond frequency measures: Can contextual embeddings capture meaning change in scientific texts?

The study investigates whether contextual embeddings can detect meaning changes in scientific terminology beyond traditional frequency counts. Using Astrophysics and NLP corpora from 2010 to 2024, the authors extract candidate terms with KeyBERT, filter for significant frequency rises, and then evaluate semantic drift via multiple embedding‑based metrics. Results show that frequency methods slightly outperform embedding metrics in aligning with expert judgments, yet embedding‑only detections (e.g., "primordial black holes") reveal critical conceptual shifts missed by frequency alone, suggesting complementary value.

By Jianying Liu (STL, BETA, CEIPI), Kim Gerdes (LISN, Qatent, STL), Jean-Marc Deltorn (CEIPI)
arXiv AI
Jun 19

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

arXiv:2605. 31393v2 Announce Type: replace-cross Abstract: Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed vocabularies typical of real-world datasets.

By Pedro Dal Bianco, Jean Paul Nunes Reinhold, Oscar Stanchi, Facundo Quiroga, Franco Ronchetti, Ulisses Brisolara Corr\^ea
arXiv Computation and Language
Sep 4

Benchmarking Machine Translation on Chinese Social Media Texts

The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.

By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao
arXiv Computation and Language
Sep 11

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

CHRONOBERG is a temporally structured corpus of English book texts covering 250 years, curated from Project Gutenberg and enriched with temporal annotations. It enables quantification of lexical semantic change via time‑sensitive Valence‑Arousal‑Dominance analysis and the creation of historically calibrated affective lexicons. Experiments show that language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts, highlighting the need for temporally aware training and evaluation pipelines.

By Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski
arXiv Computation and Language
Sep 22

Cross-Dialect NER for Bangla Regional Dialects Using Leave-One-Dialect-Out Cross-Validation and Explainable AI

The paper introduces a cross-dialect Named Entity Recognition (NER) framework for Bangla, leveraging the ANCHOLIK-NER dataset that covers five major regional dialects. Using a Leave-One-Dialect-Out Cross-Validation strategy, eight transformer-based models were evaluated, with Multilingual-E5 Large achieving the best performance (F1 up to 97.26% on Mymensingh, 82.38% on Chattogram). Local Interpretable Model-agnostic Explanations (LIME) revealed that the models rely mainly on the surface form of entity words rather than surrounding context, suggesting a direction for future improvement.

By Shamim Rahim Refat, Faika Fairuj Preotee, Shuvashis Sarker, Shifat Islam, Bidyarthi Paul, Mohammad Ashraful Hoque