arXiv Computation and Language By Danting Zhang, Bei Peng, Robert Loftin

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

Read the original on arXiv Computation and Language →

The paper introduces DECO, a diagnostic framework that factorises content into independent moderation criteria, allowing controlled evaluation of large language models (LLMs) at the criterion level. Using pairwise evaluation across four datasets and four LLMs, the authors find that high aggregate benchmark scores can mask significant failures when decisions hinge on specific content aspects required by individual criteria. The study underscores that aggregated labels do not guarantee reliable criterion-conditioned performance, highlighting the need for evaluation methods that explicitly assess this behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Sep 3

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

The paper introduces DECO, a diagnostic tool that factorises content into independent criteria for evaluating large language models (LLMs) on content moderation tasks. Using DECO and pairwise evaluation across four datasets and four LLMs, the authors find that high benchmark scores can mask significant failures at the criterion level, especially when decisions hinge on specific content aspects rather than overall harmfulness. The study underscores that aggregated label performance does not guarantee reliable criterion-conditioned evaluation, calling for new methods that explicitly assess this behavior.

arXiv Computation and Language
Sep 10

Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

arXiv:2609.10410v1 Announce Type: new Abstract: The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models po...

By Ayan Majumdar, Shounak Paul, Pushpdeep Singh, Ines Abdelaziz, Sayeh Jarollahi, Seungeon Lee, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash
arXiv AI
Jun 16

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.

By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang