arXiv AI

slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek

arXiv:2607. 21255v1 Announce Type: cross Abstract: Slang is a central component of everyday language, reflecting linguistic creativity, social identity, and cultural change, yet its dy- namic and non-standard nature makes it difficult to model computationally.

arXiv Computation and Language
Sep 23

From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits

The paper investigates how internet slang spreads across Reddit communities by combining social network analysis with linguistic context. Using large language models as scalable annotators, the authors create a benchmark for detecting slang usage and then model its adoption and diffusion. Findings reveal that users with higher bridging capital promote slang spread, while those with higher bonding capital hinder it, and that broader contextual usage delays new user adoption.

By Xiaoning Wang, Ted Underwood, Zhewei Sun
arXiv Computation and Language
Sep 15

Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection

arXiv:2606.27314v2 Announce Type: replace Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive m...

By Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu
arXiv AI
Sep 18

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

KoNeoBench is a curated dataset designed to evaluate large language models’ understanding of Korean neologisms. It contains 1,785 recently attested Korean words from online news since 2020, each accompanied by usage examples, word‑formation analyses, and dictionary‑style definitions. The authors define four evaluation tasks, report results from recent models and a human baseline, and find that current LLMs struggle with recovering source components, distinguishing semantic categories, and generating accurate definitions.

By Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam
arXiv Computation and Language
Sep 4

Benchmarking Machine Translation on Chinese Social Media Texts

The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.

By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao
arXiv Computation and Language
Sep 22

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

The BD-LSC dataset introduces a bi‑directional lexical semantic change benchmark that tracks sense gain, loss, and stability across three time periods, while the ST‑WSD dataset offers fine‑grained, instance‑level sense annotations for words that blend slang and standard usage. These resources enable systematic evaluation of diverse models—including unsupervised clustering, supervised learning, transformer‑based approaches, and large language models—on tasks such as exact sense matching and multi‑label accuracy. The evaluation shows that few‑shot GPT‑4o performs best overall, yet all systems struggle with rare slang senses, highlighting a key open challenge in the field.

By Afnan Aloraini, Riza Batista-Navarro
arXiv Computation and Language
Sep 24

The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale

The Illinois Social Attitudes Aggregate Corpus (ISAAC) is an open, modular corpus comprising over 527 million English‑language Reddit posts from 2007 to 2023, curated for relevance to six social group distinctions—race, sexuality, age, ability, body weight, and skin tone. A multi‑step, human‑audited filtering pipeline keeps irrelevant content below 10% overall and per group, while each post receives algorithmic annotations of user home region and a suite of validated semantic labels such as moralization, sentiment, emotion, and linguistic generalization. ISAAC’s publicly available, reproducible pipeline enables cross‑category comparisons, high‑precision tracking of long‑term temporal shifts, and spatial mapping of public opinion and policy outcomes, and can be extended to new platforms, languages, and social categories via both point‑and‑click and programmatic interfaces.

By Babak Hemmatian, Sarah Hadjarab, Jessica Chen, Benedek Kurdi
arXiv Computation and Language
Sep 17

Which Demographics do LLMs Default to During Annotation?

The paper investigates which demographic attributes large language models (LLMs) default to when annotating text without explicit demographic cues. By comparing non‑demographic, placebo‑conditioned, and demographic‑conditioned prompts on politeness and offensiveness tasks in the POPQUORN dataset, the authors find that LLMs exhibit notable gender, race, and age influences in their annotations. This contrasts with earlier studies that reported no such effects, highlighting the importance of considering demographic bias in LLM‑based annotation workflows.

By Johannes Sch\"afer, Aidan Combs, Christopher Bagdon, Jiahui Li, Nadine Probol, Lynn Greschner, Sean Papay, Yarik Menchaca Resendiz, Aswathy Velutharambath, Amelie W\"uhrl, Sabine Weber, Roman Klinger
arXiv Computation and Language
Sep 25

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors compare seven LLMs to human annotations. They find that models agree more with each other than with humans, show systematic neutral bias in categorical predictions, and that alignment varies with explicit linguistic cues and rapport‑building strategies.

By Rong Wang, Kun Sun, Yadong Guo