arXiv AI

VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

arXiv:2608. 15600v1 Announce Type: new Abstract: The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text.

arXiv AI
Jun 16

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.

By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang
arXiv Computation and Language
Sep 4

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

The paper introduces DECO, a diagnostic framework that factorises content into independent moderation criteria, allowing controlled evaluation of large language models (LLMs) at the criterion level. Using pairwise evaluation across four datasets and four LLMs, the authors find that high aggregate benchmark scores can mask significant failures when decisions hinge on specific content aspects required by individual criteria. The study underscores that aggregated labels do not guarantee reliable criterion-conditioned performance, highlighting the need for evaluation methods that explicitly assess this behavior.

By Danting Zhang, Bei Peng, Robert Loftin
arXiv Computation and Language
Sep 1

When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech

The paper introduces WSF-ARG+, a new dataset that pairs hate speech with check‑worthiness annotations, and presents an LLM‑in‑the‑loop framework to streamline the annotation process. Experiments with 12 open‑weight large language models demonstrate that the framework cuts human effort while maintaining annotation quality. The study also shows that incorporating check‑worthiness labels improves hate‑speech detection performance, boosting macro‑F1 scores for large models by up to 0.213 and averaging 0.154 across models.

By Nicol\'as Benjam\'in Ocampo, Tommaso Caselli, Davide Ceolin
Hugging Face Trending Papers
Sep 3

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

The paper introduces DECO, a diagnostic tool that factorises content into independent criteria for evaluating large language models (LLMs) on content moderation tasks. Using DECO and pairwise evaluation across four datasets and four LLMs, the authors find that high benchmark scores can mask significant failures at the criterion level, especially when decisions hinge on specific content aspects rather than overall harmfulness. The study underscores that aggregated label performance does not guarantee reliable criterion-conditioned evaluation, calling for new methods that explicitly assess this behavior.

arXiv AI
Sep 21

CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

CIBuzzBench is a new benchmark that tests cross‑lingual understanding of Chinese internet buzzwords by providing 3,001 buzzwords with English explanations, equivalents, category labels, and harmfulness annotations. The benchmark defines three tasks—Meaning Explanation, Cross‑lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection—to evaluate how well large language models translate and interpret these culturally nuanced terms. Experiments show that current state‑of‑the‑art LLMs still struggle with fine‑grained non‑literal meanings, robust matching, and accurate harmfulness detection across languages.

By Yifan Wang, Junyu Lu, Qifan Wang, Shun Zhang, Chaozhuo Li, Jiahao Liu, Zhijun Cao, Lingbin Bu, Fanliang Bu
arXiv Computation and Language
Aug 27

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

The paper presents an instruction‑tuned large language model (LLM) based on Qwen3 that is fine‑tuned for hate speech mitigation by unifying 36 English hate speech datasets. The authors show that this generalist LLM achieves state‑of‑the‑art performance on in‑domain benchmarks and delivers significant gains in cross‑domain and cross‑lingual generalization, outperforming specialist encoder‑based classifiers.

By Lukas Edman, Daryna Dementieva, Alexander Fraser
arXiv AI
Sep 2

Validity-Aware Jailbreak Evaluation for Large Language Models

The paper introduces SEAV, a verification‑centric framework for evaluating jailbreak attempts against large language models. SEAV decomposes responses into ordered steps and checks both validity and correctness using LLM‑as‑a‑judge and retrieval‑grounded verification. The method reduces false positives by 14.9 percentage points on a strategic‑dishonesty diagnostic and reclassifies 22.1–51.0% of previously successful jailbreaks as invalid across multiple benchmarks.

By Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran