arXiv Computation and Language
Sep 23

From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits

The paper investigates how internet slang spreads across Reddit communities by combining social network analysis with linguistic context. Using large language models as scalable annotators, the authors create a benchmark for detecting slang usage and then model its adoption and diffusion. Findings reveal that users with higher bridging capital promote slang spread, while those with higher bonding capital hinder it, and that broader contextual usage delays new user adoption.

By Xiaoning Wang, Ted Underwood, Zhewei Sun
arXiv Computation and Language
Sep 22

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

The BD-LSC dataset introduces a bi‑directional lexical semantic change benchmark that tracks sense gain, loss, and stability across three time periods, while the ST‑WSD dataset offers fine‑grained, instance‑level sense annotations for words that blend slang and standard usage. These resources enable systematic evaluation of diverse models—including unsupervised clustering, supervised learning, transformer‑based approaches, and large language models—on tasks such as exact sense matching and multi‑label accuracy. The evaluation shows that few‑shot GPT‑4o performs best overall, yet all systems struggle with rare slang senses, highlighting a key open challenge in the field.

By Afnan Aloraini, Riza Batista-Navarro
arXiv Computation and Language
Sep 4

Benchmarking Machine Translation on Chinese Social Media Texts

The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.

By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao