arXiv AI

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

The study evaluates large language models for assessing suicide risk in Arabic crisis helpline calls, comparing Arabic and English models. Using de‑identified transcripts from Lebanon’s National Lifeline, the researchers fine‑tuned instruction‑tuned LLMs and transformer encoders, achieving a macro‑F1 of 81.19 and ROC‑AUC of 90.61 for high‑risk calls in Arabic, and 85.00/92.59 in English. The results show that high‑risk calls are more distinguishable than at‑risk calls, and translating to English does not degrade performance, indicating potential for operator‑facing tools.

arXiv Computation and Language
2d ago

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

The paper introduces a framework to evaluate and diagnose the robustness of low‑resource multilingual text‑to‑speech systems when faced with complex text inputs such as numbers, dates, named entities, long sentences, code‑switched expressions, and punctuation structures. It assesses robustness across content consistency, language consistency, and generation stability, and proposes automatic metrics (character error rate, language ID accuracy, duration abnormal rate) along with a lightweight Text Risk Score (TRS) that predicts synthesis risk from interpretable text features. Experiments on Thai, Vietnamese, Swahili, and Indonesian TTS systems reveal distinct failure patterns and show that TRS correlates positively with content and duration errors, offering a low‑cost pre‑synthesis risk indicator.

By Tianlun Zuo, Ziyu Zhang, Tingzhi Mao, Zhonghua Fu, Lei Xie
arXiv AI
Sep 1

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.

By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
arXiv Computation and Language
2d ago

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.

By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi