arXiv Computation and Language
2d ago

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

The Eloquence team presents three methods for the Interspeech 2026 MLC‑SLM Task 2, a multilingual MCQA challenge covering 21 languages. They fine‑tune Voxtral‑Mini‑3B with LoRA and data augmentation, achieving 0.72 macro‑accuracy; they use multimodal in‑context learning on Voxtral‑24B to correct label bias, reaching 0.81; and they deploy a training‑free retrieval system with a voice‑anchored memory, scoring 0.68. All approaches surpass the official baseline.

By Jordi Luque, Lorenzo Concina, Marco Matassoni, Alessio Brutti, Filippo Vella
arXiv AI
Aug 26

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

EXAM$^2$ is a new benchmark for audio understanding that covers six languages and multiple modalities—speech, sound, music, mixed-audio, and visual images—providing 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations. It evaluates large audio language models (LALMs) and multimodal large language models (LLMs), revealing significant gaps in multilingual and cross‑modal performance. The authors also introduce Gemma3n-EXAM$^2$, a lightweight fusion model that improves multilingual results by up to 12.4% and multimodal results by 21.7% over a strong baseline.

By Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
arXiv Computation and Language
Aug 28

SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models

SPAR-K is a scheduled periodic alternating early‑exit framework for interleaved spoken language models that reduces decoding depth for speech tokens while maintaining quality. It lets most speech positions exit at a fixed intermediate layer and inserts periodic full‑depth refresh steps to counter distribution shift. Experiments on Step‑Audio‑2‑mini and GLM‑4‑Voice show up to 11 % depth reduction with less than 0.82 % drop in question‑answering accuracy and negligible impact on MOS and WER.

By Hsiao-Ying Huang, Cheng-Han Chiang, Hung-yi Lee