LLMs in the Real World: Evaluating "AI" in Emergency Contexts
arXiv:2607. 00019v1 Announce Type: cross Abstract: This paper offers a call to action.
arXiv:2602. 13452v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication.
arXiv:2607. 00019v1 Announce Type: cross Abstract: This paper offers a call to action.
Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis...
arXiv:2609.16997v1 Announce Type: new Abstract: Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapid...
arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
arXiv:2602.14488v3 Announce Type: replace-cross Abstract: IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensiv...
Effective crisis response requires spatially grounded communication that bridges linguistic guidance of civilians with the physical environment, accounting for structural bottlenecks, evolving threats, and agent-specific contexts. Yet, current NLP research in crisis communication remains mainly limited to static, text-only classification settings, overlooking the critical communicative role of AI operators in dynamic, embodied scenarios.
arXiv:2608. 14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages.
The paper introduces a three‑layer checklist-and-judge framework to evaluate interpreter agents that mediate live conversation across languages. It assesses semantic, pragmatic, and cultural‑social dimensions—naturalness, intent, and social appropriateness—rather than just fidelity, in both single‑turn and multi‑turn settings. Extensive validation shows that conventional MT metrics miss failures in stronger interpreters, and that context, structured instructions, and cultural cues influence communicative success.
arXiv:2601.02933v4 Announce Type: replace Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is n...
The paper investigates how machine translation can be tailored to specific audiences and intents, a capability enabled by large language models (LLMs). By systematically evaluating purpose-driven MT across 50 languages, 5 model sizes, and 8 text domains, the authors find that explicit instructions significantly improve translation adaptiveness, especially for informal domains, larger models, and higher-resource languages. They also show that traditional MT metrics often penalize adapted translations and that models can self-generate useful instructions from context, closing a large portion of the adaptiveness gap.
The paper presents an LLM-based framework for automatically classifying crisis levels in psychological support hotlines, addressing variability in human judgments and staffing constraints. It introduces a paralinguistic injection method that embeds non‑verbal emotional cues into transcripts, allowing the model to consider acoustic nuances. A reasoning‑enhanced training strategy encourages the model to produce diagnostic reasoning chains, which regularizes and improves classification, achieving a macro F1‑score of 0.802 and accuracy of 0.805 in 5‑fold cross‑validation.
The paper introduces a framework to evaluate and diagnose the robustness of low‑resource multilingual text‑to‑speech systems when faced with complex text inputs such as numbers, dates, named entities, long sentences, code‑switched expressions, and punctuation structures. It assesses robustness across content consistency, language consistency, and generation stability, and proposes automatic metrics (character error rate, language ID accuracy, duration abnormal rate) along with a lightweight Text Risk Score (TRS) that predicts synthesis risk from interpretable text features. Experiments on Thai, Vietnamese, Swahili, and Indonesian TTS systems reveal distinct failure patterns and show that TRS correlates positively with content and duration errors, offering a low‑cost pre‑synthesis risk indicator.