arXiv Computation and Language By Chaewan Chun, Delvin Ce Zhang, Dongwon Lee

Context-Aware Multimodal Claim Verification in Spoken Dialogues

Read the original on arXiv Computation and Language →

The paper introduces MAD2, a synthetic benchmark of 1,000 two‑speaker dialogues with about 10 hours of audio and 1,230 check‑worthy sentence annotations for spoken claim verification. It proposes a calibrated multimodal fusion approach that combines a context‑aware audio encoder with a dialogue‑aware text model. Experiments show that adding dialogue context improves verification performance, though the gains differ across scenarios, and that fusion offers the largest advantage when full‑dialogue context is available, though it does not consistently outperform text alone.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 28

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.

By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
arXiv Computation and Language
Sep 21

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

The paper introduces Omni Demand Understanding (ODU), a benchmark designed to test whether multimodal models can infer a user's underlying demand from complex audio‑visual interactions. ODU requires models to detect the presence of a demand and infer intent using multimodal and conversational context, evaluated across single‑turn and multi‑turn scenarios. The authors built ODU‑Bench through a taxonomy‑guided approach, agentic video generation, and human‑recorded interactions, and found that even top models like Gemini 3.1 Pro recover only 44.7% of key information, with many models exhibiting high false‑trigger rates.

By Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang, Ziyang Ma, Ruiyang Xu, Meng Gao, Yinsong Yan, Ling Wang, Hui Wang, Wen Huang, Yiheng Chen, Guanrou Yang, Qiuqiang Kong, Jin Xu, Xie Chen
arXiv Computation and Language
Sep 7

TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio

TRILOGUE is a new trilingual benchmark for spoken dialogue fact‑checking, covering English, Russian, and Kazakh. It includes almost 12,000 dialogues, 187,000 turns, and 390 hours of paired audio with ASR transcripts and word‑level timestamps, as well as nearly 5,000 human‑recorded Russian and Kazakh files. The dataset supports tasks such as claim check‑worthiness detection, evidence retrieval, and claim verification under various input conditions, and baseline experiments reveal challenges with ASR errors and cross‑lingual transfer, especially for Kazakh.

By Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi, Mahjabin Nahar, Delvin Ce Zhang, Dongwon Lee
arXiv AI
5d ago

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

The paper introduces VeriSpeak, a benchmark of 3,879 spoken claims for evaluating fact verification in Large Audio Language Models (LALMs). It shows a clear modality gap: models that verify written claims well often fail on spoken versions, and retrieval alone offers limited improvement. Combining retrieval with explicit reasoning yields the best performance, reaching 86.1% accuracy and demonstrating the need for grounded reasoning over retrieved evidence in speech misinformation detection.

By Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen