arXiv AI By Tomer Atia, Yehudit Aperstein, Alexander Apartsin

SeaAlert: Robust Severity Classification and LLM-Based Information Extraction for Noisy Maritime Distress Communications

Read the original on arXiv AI →

arXiv:2604. 14163v2 Announce Type: replace-cross Abstract: Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergencies at sea.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Speech-based Psychological Crisis Assessment using LLMs

The paper presents an LLM-based framework for automatically classifying crisis levels in psychological support hotlines, addressing variability in human judgments and staffing constraints. It introduces a paralinguistic injection method that embeds non‑verbal emotional cues into transcripts, allowing the model to consider acoustic nuances. A reasoning‑enhanced training strategy encourages the model to produce diagnostic reasoning chains, which regularizes and improves classification, achieving a macro F1‑score of 0.802 and accuracy of 0.805 in 5‑fold cross‑validation.

By Terumi Chiba, Yang Luo, Ziyun Cui, Yongsheng Tong, Chao Zhang
arXiv AI
Sep 2

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

The study evaluates large language models for assessing suicide risk in Arabic crisis helpline calls, comparing Arabic and English models. Using de‑identified transcripts from Lebanon’s National Lifeline, the researchers fine‑tuned instruction‑tuned LLMs and transformer encoders, achieving a macro‑F1 of 81.19 and ROC‑AUC of 90.61 for high‑risk calls in Arabic, and 85.00/92.59 in English. The results show that high‑risk calls are more distinguishable than at‑risk calls, and translating to English does not degrade performance, indicating potential for operator‑facing tools.

By Linhai Ma, Rita El Hachem, Mahatab El Hajj, Lilian Ghandour, Samah Fodeh
arXiv Computation and Language
Sep 11

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

The paper introduces a framework to evaluate and diagnose the robustness of low‑resource multilingual text‑to‑speech systems when faced with complex text inputs such as numbers, dates, named entities, long sentences, code‑switched expressions, and punctuation structures. It assesses robustness across content consistency, language consistency, and generation stability, and proposes automatic metrics (character error rate, language ID accuracy, duration abnormal rate) along with a lightweight Text Risk Score (TRS) that predicts synthesis risk from interpretable text features. Experiments on Thai, Vietnamese, Swahili, and Indonesian TTS systems reveal distinct failure patterns and show that TRS correlates positively with content and duration errors, offering a low‑cost pre‑synthesis risk indicator.

By Tianlun Zuo, Ziyu Zhang, Tingzhi Mao, Zhonghua Fu, Lei Xie
arXiv Machine Learning
Sep 22

Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs

The paper introduces MMSAFE, a multi-layer framework designed to identify safety-degrading data in multilingual large language models. It shows that safety signals are distributed across multiple layers and only partially shared across languages, unlike the single-layer assumption used in monolingual settings. Experiments demonstrate that MMSAFE reduces harmful-response rates by 60% compared to random filtering and outperforms the best single-layer baseline across various models, languages, and safety benchmarks.

By Jiakun Li, Guowei Song, Sijia Li, Xingwei He, Hongzheng Chai, Yuan Yuan