arXiv:2607. 15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch.
By Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana
Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention.
arXiv:2606.27314v2 Announce Type: replace
Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive m...
By Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu
arXiv:2605.28013v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vi...
By Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon, Dujin Lee, HyunBeom Cho, Jaewon Lee, Wonhyuk Lee, Youngchol Kim, JeongYeop Kim, Donghyun Kim
AraDetox is a newly released multi-dialect Arabic detoxification dataset containing 10,500 harmful social‑media posts and 84,000 detoxified rewrites generated by GPT‑5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. Human evaluation and automatic analyses confirm that the rewrites effectively remove harmful language while preserving meaning, lexical change, and dialectal style. The dataset is publicly available to support future research in Arabic detoxification, safe text generation, and multi‑dialect NLP.
By Mo El-Haj
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark covers ten safety-critical content categories, six persuasive strategies, and four languages—Hindi, Bengali, Marathi, and Punjabi—producing 7,200 adversarial prompts. Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, highlighting gaps in current English-centric safety assessments.
arXiv:2605. 14152v2 Announce Type: replace-cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual safety is mostly assessed through translation-only benchmarks that preserve the underlying scenario, leaving how language and geopolitical context interact largely unexamined beyond a few language pairs.
By Michael S. Lee, Yash Maurya, Drew Rein, Bert Herring, Jonathan Nguyen, Kyungho Song, Udari Madhushani Sehwag, Jiyeon Cho, Kaustubh Deshpande, Yeongkyun Jang, Jiyeon Joo, Minn Seok Choi, Evi Fuelle, Christina Q. Knight, Joseph Brandifino, Max Fenkell
arXiv:2607. 23175v1 Announce Type: cross Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent.
By Rares A. C. Diaconescu, Iulia Slanina, Alina Florea, Andrei B. Trache, Miruna E. Coroi, Anne Arzberger, Jie Yang, Enrico Liscio
arXiv:2510. 10271v2 Announce Type: replace-cross Abstract: Unlike regular tokens derived from existing text corpora, special tokens are artificially created to annotate structured conversations during the fine-tuning process of Large Language Models (LLMs).
By Wentian Zhu, Zhen Xiang, Wei Niu, Le Guan
arXiv:2507. 10177v2 Announce Type: replace-cross Abstract: Although Large Language Models (LLMs) have demonstrated significant advancements in natural language processing tasks, their effectiveness in the classification and transformation of abusive text into non-abusive versions remains an area for exploration.
By Rohitash Chandra, Jiyong Choi, Jayesh Sonawane
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark includes 7,200 adversarial prompts covering ten safety-critical content categories, six persuasive strategies, and four languages (Hindi, Bengali, Marathi, Punjabi). Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, revealing that current English-centric safety evaluations miss important multilingual vulnerabilities.
By Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal
arXiv:2606. 15396v1 Announce Type: cross Abstract: Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns.
By Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao, Jingru Zeng, Xiaochen Yang, Tianyi Zhang, Xiaoxiao Ma, Jiawei Kong, Hao Wu, Bin Chen, Shu-Tao Xia, Min Zhang