arXiv Computation and Language
3d ago

Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection

The paper introduces a diagnostic tool for distinguishing the use of misogynistic slurs from their mention in counter‑speech within code‑mixed Hinglish. It identifies evaluation artifacts in existing corpora, releases a 416‑item minimal‑pair contrast set that decorrelates slur presence and gendered register from labels, and proposes a pair‑consistency metric to assess model performance. Experiments show that even strong baselines struggle to consistently label counter‑speech pairs, while a large language model achieves perfect scores, indicating the benchmark measures genuine capability rather than exploitation of artifacts.

By Ashanvi Yadav, Shubham Bhardwaj
arXiv AI
Aug 24

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

The paper identifies a vulnerability in large language models where harmful intent can be hidden within benign narratives, a phenomenon termed Semantic Camouflage. By examining latent activation patterns across several small language model families, the authors discover an "Intent Horizon"—a layer depth where harmful intent representations collapse. They propose Latent Intent Verification (LIV), a lightweight probing defense that detects harmful intent in early layers and outperforms existing guardrails on the PKU-SafeRLHF dataset.

By Md. Hasib Ur Rahman
arXiv AI
4d ago

CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

CIBuzzBench is a new benchmark that tests cross‑lingual understanding of Chinese internet buzzwords by providing 3,001 buzzwords with English explanations, equivalents, category labels, and harmfulness annotations. The benchmark defines three tasks—Meaning Explanation, Cross‑lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection—to evaluate how well large language models translate and interpret these culturally nuanced terms. Experiments show that current state‑of‑the‑art LLMs still struggle with fine‑grained non‑literal meanings, robust matching, and accurate harmfulness detection across languages.

By Yifan Wang, Junyu Lu, Qifan Wang, Shun Zhang, Chaozhuo Li, Jiahao Liu, Zhijun Cao, Lingbin Bu, Fanliang Bu
arXiv AI
Aug 3

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

arXiv:2607. 28862v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage.

By Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu
arXiv AI
Aug 11

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

arXiv:2608. 09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale.

By Kevin Thomas, Milosz Kasprzyk, Reuel C Igbokwe Onuigbo, Elliott Pert, Cameron Tovey, Jo\~ao A. Leite, Olesya Razuvayevskaya, Carolina Scarton