The paper introduces DECO, a diagnostic framework that factorises content into independent moderation criteria, allowing controlled evaluation of large language models (LLMs) at the criterion level. Using pairwise evaluation across four datasets and four LLMs, the authors find that high aggregate benchmark scores can mask significant failures when decisions hinge on specific content aspects required by individual criteria. The study underscores that aggregated labels do not guarantee reliable criterion-conditioned performance, highlighting the need for evaluation methods that explicitly assess this behavior.
By Danting Zhang, Bei Peng, Robert Loftin
The paper introduces DECO, a diagnostic tool that factorises content into independent criteria for evaluating large language models (LLMs) on content moderation tasks. Using DECO and pairwise evaluation across four datasets and four LLMs, the authors find that high benchmark scores can mask significant failures at the criterion level, especially when decisions hinge on specific content aspects rather than overall harmfulness. The study underscores that aggregated label performance does not guarantee reliable criterion-conditioned evaluation, calling for new methods that explicitly assess this behavior.
arXiv:2609.22094v1 Announce Type: cross
Abstract: Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining fo...
By Zeeshan Ahmed, Yang Qin, Hanqing Huang
arXiv:2604. 06205v2 Announce Type: replace-cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types.
By Shutong Zhang, Dylan Zhou, Yinxiao Liu, Yang Yang, Huiwen Luo, Wenfei Zou
arXiv:2606. 05748v1 Announce Type: cross Abstract: Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement.
By Kejuan Yang, Yizhuo Zhang, Mingyuan Du, Yue Zhang, Dixin Zheng, Kaili Zhao, Yang Xiao, Hanzhong Liang, Kenan Xiao
The paper introduces a training‑free approach to detect policy violations in large language models by treating the task as an out‑of‑distribution problem in the model’s activation space. It uses whitening‑inspired techniques to compute policy‑violation scores directly from normalized hidden activations, requiring only the policy text and a few illustrative examples. Experiments on several LLMs and policy benchmarks show the method achieves up to 86.0% F1, outperforming fine‑tuned and LLM‑as‑a‑judge baselines while being computationally lightweight.
By Oren Rachmil, Avishag Shapira, Roy Betser, Omer Hofman, Itay Gershon, Asaf Shabtai, Yuval Elovici, Roman Vainshtein
arXiv:2608. 19207v1 Announce Type: new Abstract: Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior.
By Juan Yeo, Geewook Kim
arXiv:2607. 20083v1 Announce Type: cross Abstract: Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models.
By Beining Wang, Weihang Su, Hongtao Tian, Hao Kong, Tao Yang, Ting Yao, Qingyi Pan, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu
The paper presents a scalable approach to harmful content moderation on social media by leveraging large language models (LLMs) for few-shot, in-context learning. Experiments across multiple LLMs show that this method outperforms proprietary baselines such as Perspective and OpenAI Moderation, as well as prior few-shot learning techniques, in detecting harmful content. The study also explores the addition of visual cues like video thumbnails to assess multimodal improvements, highlighting the advantages of LLM-based moderation for dynamic and large-scale content filtering.
By Akash Bonagiri, Lucen Li, Rajvardhan Oak, Zeerak Babar, Magdalena Wojcieszak, Anshuman Chhabra
The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.
By Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar, Vabuk Pahari, Abhijnan Chakraborty, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash
The paper introduces KoNA, a benchmark designed to evaluate selective non‑compliance in vision‑language models (VLMs) across five categories—False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety. KoNA tests both query‑level and component‑level non‑compliance using paired single and compound queries, revealing that many VLMs struggle to refuse, correct, or abstain appropriately, especially when selective non‑compliance is required. Fine‑tuning VLMs on KoNA examples improves non‑compliance accuracy while preserving performance on fully answerable tasks, indicating that models can better distinguish answerable components from those needing non‑compliance.
By Minji Kim, Jihyoung Jang, Hyounghun Kim
arXiv:2503. 06573v3 Announce Type: replace-cross Abstract: Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge.
By Gili Lior, Asaf Yehudai, Ariel Gera, Liat Ein-Dor