arXiv Machine Learning By Julia Kruk, Sanchita Porwal, Amitrajit Bhattacharjee, Mansi Phute

Community-Specific Slang and Entity Detection via Semantic Shift in Fine-Tuned Language Models

Read the original on arXiv Machine Learning →

arXiv:2606. 07522v1 Announce Type: cross Abstract: We propose an unsupervised method of resolving slang, unique entities, and folklore from online communities by isolating words in the lexicon that have the highest magnitude of semantic shift.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 23

From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits

The paper investigates how internet slang spreads across Reddit communities by combining social network analysis with linguistic context. Using large language models as scalable annotators, the authors create a benchmark for detecting slang usage and then model its adoption and diffusion. Findings reveal that users with higher bridging capital promote slang spread, while those with higher bonding capital hinder it, and that broader contextual usage delays new user adoption.

By Xiaoning Wang, Ted Underwood, Zhewei Sun
arXiv Computation and Language
Sep 22

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

The BD-LSC dataset introduces a bi‑directional lexical semantic change benchmark that tracks sense gain, loss, and stability across three time periods, while the ST‑WSD dataset offers fine‑grained, instance‑level sense annotations for words that blend slang and standard usage. These resources enable systematic evaluation of diverse models—including unsupervised clustering, supervised learning, transformer‑based approaches, and large language models—on tasks such as exact sense matching and multi‑label accuracy. The evaluation shows that few‑shot GPT‑4o performs best overall, yet all systems struggle with rare slang senses, highlighting a key open challenge in the field.

By Afnan Aloraini, Riza Batista-Navarro
arXiv Computation and Language
Sep 22

Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments

The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.

By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte