arXiv:2403. 18957v3 Announce Type: replace-cross Abstract: Online user generated content games (UGCGs) are increasingly popular among children and adolescents for social interaction and more creative online entertainment.
By Keyan Guo, Ayush Utkarsh, Wenbo Ding, Isabelle Ondracek, Ziming Zhao, Guo Freeman, Nishant Vishwamitra, Hongxin Hu
The paper introduces COMIC, a reference‑aware safety gate designed for multimodal large language models (MLLMs). COMIC detects the operation requested by a user, identifies visual targets through OCR and open‑vocabulary proposals, and evaluates safety on explicit operation‑target pairs, using max‑risk aggregation and quality‑aware routing to decide whether to allow or block a request. Experiments on several open‑source MLLMs and jailbreak benchmarks show that COMIC improves robustness while maintaining benign utility and efficiency.
By Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu, Chenyu Wang, Honghui Xu
Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization.
The paper introduces ProKDA, a progressive knowledge-to-decision alignment framework for explainable hateful meme detection. ProKDA separates explanation generation and label prediction into three sequential training stages—background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment—reducing task interference. Experiments on three public benchmarks demonstrate that ProKDA achieves state‑of‑the‑art detection performance while providing accurate, evidence‑supported explanations for moderation decisions.
By Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia
arXiv:2609.14677v1 Announce Type: new
Abstract: Large language models (LLMs) have changed the way people engage with stories. Drawing on public chatbot logs, we can see that when users generate stori...
By Advait Deshmukh, Nora Benedict, Melanie Walsh, Maria Antoniak
DiSCO is a zero‑shot, black‑box defense for text‑to‑image models that operates solely at the prompt level. It expands prompts with a distribution‑guided suffix using beam search and contrastive scoring against safe and unsafe image pools generated by the target model, iteratively refining until safe content is produced. The method improves safety on the I2P benchmark under various red‑teaming attacks, reducing attack success rates by 37.7% and 25.13% while preserving semantic fidelity and image coherence.
By Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
arXiv:2608.23152v1 Announce Type: new
Abstract: Counterspeech effectively neutralizes the impact of online hate. Although prior work explores automated counterspeech generation, it largely emphasizes...
By Sujoy Nath, Aswini Kumar, Tanmoy Chakraborty
arXiv:2606. 09700v1 Announce Type: cross Abstract: Large language model (LLM)-powered content moderation systems have become a critical defense against harmful online content.
By Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang, Meisam Mohammady, Doowon Kim, Yuan Hong
InGuard introduces an inner guardrail for text-to-image generation that operates within the model’s own representations, avoiding external classifiers. It grades prompts using the text encoder’s embeddings, modifies risky embeddings with SAGE to produce safe images, and employs a latent detector to halt generation early. Evaluated on the RevGen Safety Benchmark, InGuard achieves a 97.9–98.8% safety rate across five open-weight models while reducing benign disturbances, model parameters, and denoising steps.
By Zeyu Wang, Xiaodan Li, Zhiwen Li, Yuefeng Chen, Hui Xue
arXiv:2512. 08724v3 Announce Type: replace Abstract: Text-to-image (TTI) diffusion models have achieved remarkable visual quality, yet they have been repeatedly shown to exhibit social biases across sensitive attributes such as gender, race and age.
By Manos Plitsis, Giorgos Bouritsas, Vassilis Katsouros, Yannis Panagakis
CollageAttack is a black‑box jailbreak that exploits cross‑modal alignment flaws in text‑to‑image models by shifting harmful semantics into the image plane. It combines context‑relevant scenes, scene‑grounded textual carriers, and spatially distributed text fragments to produce images that reveal hidden harmful meaning. Experiments on both open‑weight and commercial models show success rates up to 86.0%, outperforming the strongest baseline by 18.5 percentage points and consistently generating more harmful outputs while preserving the source intent.
By Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren
arXiv:2606. 31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models.
By Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong, Evan Frick, I-Hung Hsu, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh