The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
The paper introduces a framework to evaluate and diagnose the robustness of low‑resource multilingual text‑to‑speech systems when faced with complex text inputs such as numbers, dates, named entities, long sentences, code‑switched expressions, and punctuation structures. It assesses robustness across content consistency, language consistency, and generation stability, and proposes automatic metrics (character error rate, language ID accuracy, duration abnormal rate) along with a lightweight Text Risk Score (TRS) that predicts synthesis risk from interpretable text features. Experiments on Thai, Vietnamese, Swahili, and Indonesian TTS systems reveal distinct failure patterns and show that TRS correlates positively with content and duration errors, offering a low‑cost pre‑synthesis risk indicator.
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
arXiv:2606.06037v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evalu...
arXiv:2608. 13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users.
arXiv:2608. 14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages.
arXiv:2606. 08451v1 Announce Type: cross Abstract: Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy.
arXiv:2608.28776v1 Announce Type: new Abstract: Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translat...
arXiv:2605. 17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages.
arXiv:2608. 20047v1 Announce Type: cross Abstract: Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements.
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark covers ten safety-critical content categories, six persuasive strategies, and four languages—Hindi, Bengali, Marathi, and Punjabi—producing 7,200 adversarial prompts. Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, highlighting gaps in current English-centric safety assessments.
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark includes 7,200 adversarial prompts covering ten safety-critical content categories, six persuasive strategies, and four languages (Hindi, Bengali, Marathi, Punjabi). Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, revealing that current English-centric safety evaluations miss important multilingual vulnerabilities.
The paper titled "Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs" highlights that current safety alignment training for large language models is predominantly English-centric, leading to failures in non‑English languages. It introduces INCLUDE, a multilingual benchmark with 2,604 prompts in six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to measure Indian‑centric socio‑cultural biases. Evaluation of ten open‑ and closed‑source LLMs shows that Bengali models exhibit the highest bias scores among open‑source models, while English shows the lowest bias in open‑source but the highest in closed‑source models.