The paper investigates how internet slang spreads across Reddit communities by combining social network analysis with linguistic context. Using large language models as scalable annotators, the authors create a benchmark for detecting slang usage and then model its adoption and diffusion. Findings reveal that users with higher bridging capital promote slang spread, while those with higher bonding capital hinder it, and that broader contextual usage delays new user adoption.
By Xiaoning Wang, Ted Underwood, Zhewei Sun
arXiv:2606. 07522v1 Announce Type: cross Abstract: We propose an unsupervised method of resolving slang, unique entities, and folklore from online communities by isolating words in the lexicon that have the highest magnitude of semantic shift.
By Julia Kruk, Sanchita Porwal, Amitrajit Bhattacharjee, Mansi Phute
arXiv:2607. 21255v1 Announce Type: cross Abstract: Slang is a central component of everyday language, reflecting linguistic creativity, social identity, and cultural change, yet its dy- namic and non-standard nature makes it difficult to model computationally.
By Panagiotis Papadakos, Katerina Papantoniou, Dimitris Plexousakis
The BD-LSC dataset introduces a bi‑directional lexical semantic change benchmark that tracks sense gain, loss, and stability across three time periods, while the ST‑WSD dataset offers fine‑grained, instance‑level sense annotations for words that blend slang and standard usage. These resources enable systematic evaluation of diverse models—including unsupervised clustering, supervised learning, transformer‑based approaches, and large language models—on tasks such as exact sense matching and multi‑label accuracy. The evaluation shows that few‑shot GPT‑4o performs best overall, yet all systems struggle with rare slang senses, highlighting a key open challenge in the field.
By Afnan Aloraini, Riza Batista-Navarro
arXiv:2510. 18908v2 Announce Type: replace-cross Abstract: Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic.
By Wangjiaxuan Xin, Shuhua Yin, Shi Chen, Yaorong Ge
The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.
By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao
arXiv:2607. 14957v1 Announce Type: new Abstract: Online firestorms are rapid collective escalations of highly negative user-generated content and may cause substantial reputational and economic damage.
By Besim Shala, Peter Mandl, Andreas Humpe, Martin H\"ausl
The Illinois Social Attitudes Aggregate Corpus (ISAAC) is an open, modular corpus comprising over 527 million English‑language Reddit posts from 2007 to 2023, curated for relevance to six social group distinctions—race, sexuality, age, ability, body weight, and skin tone. A multi‑step, human‑audited filtering pipeline keeps irrelevant content below 10% overall and per group, while each post receives algorithmic annotations of user home region and a suite of validated semantic labels such as moralization, sentiment, emotion, and linguistic generalization. ISAAC’s publicly available, reproducible pipeline enables cross‑category comparisons, high‑precision tracking of long‑term temporal shifts, and spatial mapping of public opinion and policy outcomes, and can be extended to new platforms, languages, and social categories via both point‑and‑click and programmatic interfaces.
By Babak Hemmatian, Sarah Hadjarab, Jessica Chen, Benedek Kurdi
The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.
By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
The paper investigates how shared community affiliations, measured via Bluesky starter packs, correlate with common ground between users. By analyzing 191,648 user pairs, it finds that lexical similarity—used as a proxy for common ground—increases monotonically with the number of shared starter packs, especially when those packs represent distinct topical communities. The study also shows that this effect is independent of network proximity, indicating that community membership is a distinct, measurable carrier of common ground.
By Sagar Kumar, Lawrence Swaminathan Xavier Prince, Julia Mendelsohn, Brooke Foucault Welles, Nicholas W. Landry
arXiv:2606. 02883v1 Announce Type: cross Abstract: Recommender systems have grown from content-organization tools into sophisticated systems that shape daily behavior.
By Amir Ghasemian, Homa Hosseinmardi, Upasana Dutta, Duncan J. Watts
The paper investigates whether large language models (LLMs) can generate synthetic cyberbullying conversations that replicate the social dynamics of real interactions. Using a comprehensive framework, the authors compare authentic dialogues with synthetic ones from GPT, Grok, and LLaMA across structural, linguistic, affective, and temporal dimensions, and conduct human evaluations of realism. Results show that while LLMs preserve high‑level interaction patterns, they systematically distort finer‑grained social phenomena, with model‑specific biases such as GPT’s suppression of harmful content and Grok’s amplification of aggression.
By Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis