arXiv AI By Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana

Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

Read the original on arXiv AI →

arXiv:2607. 15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs

The paper introduces MMSAFE, a multi-layer framework designed to identify safety-degrading data in multilingual large language models. It shows that safety signals are distributed across multiple layers and only partially shared across languages, unlike the single-layer assumption used in monolingual settings. Experiments demonstrate that MMSAFE reduces harmful-response rates by 60% compared to random filtering and outperforms the best single-layer baseline across various models, languages, and safety benchmarks.

By Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan