arXiv AI

DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset

arXiv:2606. 00012v1 Announce Type: cross Abstract: Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations.

arXiv AI
Jul 3

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

arXiv:2607. 02504v1 Announce Type: cross Abstract: Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character.

By Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian
arXiv Computation and Language
Sep 18

The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation

The Public Discourse Corpus (PDC) is the first dataset of public‑figure interview speech annotated for affective valence and epistemic modality. It contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words). A key methodological contribution is Target Speaker Participation (TSP), a five‑category annotation taxonomy with documented inter‑annotator reliability (κ = 0.616), and an audio‑first diarization pipeline that separates target‑speaker turns from interviewer and third‑party speech. The corpus, annotation tools, validation sample, and processing pipeline are released as open source.

By Bo Chen
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv Machine Learning
Jul 28

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

arXiv:2607. 23242v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms.

By Sahil Deepak Gawande, Mayank Singh
arXiv Computation and Language
Sep 24

Context-Aware Multimodal Claim Verification in Spoken Dialogues

The paper introduces MAD2, a synthetic benchmark of 1,000 two‑speaker dialogues with about 10 hours of audio and 1,230 check‑worthy sentence annotations for spoken claim verification. It proposes a calibrated multimodal fusion approach that combines a context‑aware audio encoder with a dialogue‑aware text model. Experiments show that adding dialogue context improves verification performance, though the gains differ across scenarios, and that fusion offers the largest advantage when full‑dialogue context is available, though it does not consistently outperform text alone.

By Chaewan Chun, Delvin Ce Zhang, Dongwon Lee
arXiv Computation and Language
Sep 3

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

TalkFa introduces a unified benchmark for Farsi dialogue generation and understanding, comprising three datasets: WIKI‑FADIAL (4.2K Wikipedia‑grounded dialogues), DAILYDIALOG‑FA (6.6K dialogues with dialogue‑act and emotion annotations), and PLAYDIAL‑FA (2.1K theatrical dialogues with sentiment labels). All dialogues are curated through multi‑stage review by native speakers, ensuring high quality. Experiments show that LoRA fine‑tuning improves generation performance with less data, while specific models excel on classification tasks, and human evaluation confirms the benchmark’s reliability.

By Neda Jamshidi, Kamyar Zeinalipour, Fahimeh Akbari, Monica Bianchini, Marco Maggini, Marco Gori
arXiv Computation and Language
Aug 27

Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting

The paper introduces a method for forecasting conversational derailment by incorporating speech act information as an auxiliary signal to enhance pragmatic representations. This approach aims to reduce lexical noise and improve generalizability, especially in low-data and cross-domain scenarios. Experiments on three datasets demonstrate performance gains over existing methods.

By Angela Yifei Yuan, Christine De Kock, Christopher Leckie