arXiv Computation and Language
Sep 11

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

The paper presents SEAR, a system for Task 2 of the second Multilingual Conversational Speech Language Model Challenge. It adapts Qwen3-Omni-30B-A3B-Instruct with a segment‑evidence‑aware pipeline that converts timestamped ASR into event spans, expands and crops them, and then generates both semantic and acoustic multiple‑choice questions using Qwen3.6-27B and Gemini 3.1 Flash‑Lite. After rigorous checks, 359,825 verified MCQs across 21 languages and accents are produced, and the system achieves 90.92% accuracy on the official evaluation set.

By Huy Hoang Le, Long-Bao Nguyen, Minh Tri Dao
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv AI
Aug 28

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.

By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
arXiv Computation and Language
Sep 11

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

The Eloquence team presents three methods for the Interspeech 2026 MLC‑SLM Task 2, a multilingual MCQA challenge covering 21 languages. They fine‑tune Voxtral‑Mini‑3B with LoRA and data augmentation, achieving 0.72 macro‑accuracy; they use multimodal in‑context learning on Voxtral‑24B to correct label bias, reaching 0.81; and they deploy a training‑free retrieval system with a voice‑anchored memory, scoring 0.68. All approaches surpass the official baseline.

By Jordi Luque, Lorenzo Concina, Marco Matassoni, Alessio Brutti, Filippo Vella