Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition
arXiv:2607. 14769v1 Announce Type: cross Abstract: Existing text summarization research has focused much on monologic information (e.
arXiv:2606. 00012v1 Announce Type: cross Abstract: Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations.
arXiv:2607. 14769v1 Announce Type: cross Abstract: Existing text summarization research has focused much on monologic information (e.
arXiv:2607. 02504v1 Announce Type: cross Abstract: Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character.
arXiv:2607. 23242v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms.
arXiv:2607. 24191v1 Announce Type: cross Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling.
arXiv:2606. 20400v1 Announce Type: new Abstract: Generating high-utility synthetic data for intent classification typically requires human-annotated seed data, which is often unavailable in fast-paced industrial settings.
arXiv:2603. 18558v2 Announce Type: replace-cross Abstract: Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows.
arXiv:2509. 09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline.
Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs).
arXiv:2605. 21182v2 Announce Type: replace-cross Abstract: Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture.
arXiv:2606. 07533v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues.
arXiv:2405. 12775v2 Announce Type: replace-cross Abstract: Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.
arXiv:2507. 19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework.