arXiv:2607. 15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch.
By Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana
Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention.
arXiv:2606.27314v2 Announce Type: replace
Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive m...
By Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu
arXiv:2605.28013v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vi...
By Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon, Dujin Lee, HyunBeom Cho, Jaewon Lee, Wonhyuk Lee, Youngchol Kim, JeongYeop Kim, Donghyun Kim
AraDetox is a newly released multi-dialect Arabic detoxification dataset containing 10,500 harmful social‑media posts and 84,000 detoxified rewrites generated by GPT‑5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. Human evaluation and automatic analyses confirm that the rewrites effectively remove harmful language while preserving meaning, lexical change, and dialectal style. The dataset is publicly available to support future research in Arabic detoxification, safe text generation, and multi‑dialect NLP.
By Mo El-Haj
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark covers ten safety-critical content categories, six persuasive strategies, and four languages—Hindi, Bengali, Marathi, and Punjabi—producing 7,200 adversarial prompts. Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, highlighting gaps in current English-centric safety assessments.