arXiv Computation and Language
Sep 16

PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark

PARSA‑Bench is the first dedicated benchmark for evaluating large audio‑language models on Persian, addressing unique challenges such as classical poetry, traditional music, and code‑switching. It comprises 16 tasks—10 of which are new—covering speech understanding, paralinguistic analysis, and culturally grounded audio reasoning. Across most tasks, text‑only baselines outperform audio‑based models, indicating that audio understanding remains the main limitation, except for Persian poetry where prosody provides additional information that audio beats text.

By Mohammad Javad Ranjbar Kalahroodi, Mohammad Amini, Parmis Bathayan, Heshaam Faili, Azadeh Shakery
arXiv Computation and Language
Aug 31

Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

This paper investigates whether prosodic features—pitch, energy, and timing—are preserved when speech is translated between languages. Using multilingual dubbing data for English‑German, English‑Spanish, and English‑French pairs, the authors conduct a fine‑grained cross‑lingual analysis to quantify similarities and differences in prosody. The study identifies inherent cross‑lingual correlations in prosodic structure and explores how linguistic and alignment factors influence these patterns.

By Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn