Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
Conversational DNA is a visual language and interactive atlas designed to explore human and AI dialogue by mapping speaker strands, communicative bases, and directed pairings. It visualizes speaker switching, response distance, and contribution length through adjustable helix geometry, and covers 151,489 episodes across eight corpora totaling 1.57 million source records. The system improves precision@5 on Molweni motif queries from 58.8% to 77.2% and demonstrates how annotation coverage affects perceived collection differences.
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
arXiv:2601. 19792v4 Announce Type: replace-cross Abstract: For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical.
Lexara-RF introduces reference‑free metrics for evaluating conversational visual analytics agents that generate visualizations and natural‑language explanations. The framework uses only the prompt, data, and model response to score outputs, applying 13 metrics derived from visualization design theory and Gricean principles as consistency, intent‑alignment, and design validity checks. In tests against a human‑rated corpus, Lexara‑RF matches reference‑based methods, outperforms surface‑similarity NLG baselines, and accurately identifies structurally grounded failures.
arXiv:2609.00802v1 Announce Type: new Abstract: Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate i...
arXiv:2606. 08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.
The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.
arXiv:2606. 09169v1 Announce Type: new Abstract: In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.
Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs).
arXiv:2608. 06975v1 Announce Type: cross Abstract: Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative.
arXiv:2607. 16208v1 Announce Type: new Abstract: Graph-grounded multimodal question answering organizes text, tables, and images in a structured evidence graph, yet end-to-end accuracy depends on which multimodal assets are ranked highly enough to enter downstream reasoning; for graph-linked images, single-vector bi-encoder similarity can discard patch- and token-level structure needed for fine-grained alignment.
Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time.
arXiv:2606. 00370v1 Announce Type: cross Abstract: Diverse genomics data, scientific questions, and analysis tasks typically demand highly specialized visualizations.