arXiv AI

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

arXiv:2603. 16859v2 Announce Type: replace Abstract: Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text.

arXiv Computation and Language
Sep 21

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

The paper introduces Omni Demand Understanding (ODU), a benchmark designed to test whether multimodal models can infer a user's underlying demand from complex audio‑visual interactions. ODU requires models to detect the presence of a demand and infer intent using multimodal and conversational context, evaluated across single‑turn and multi‑turn scenarios. The authors built ODU‑Bench through a taxonomy‑guided approach, agentic video generation, and human‑recorded interactions, and found that even top models like Gemini 3.1 Pro recover only 44.7% of key information, with many models exhibiting high false‑trigger rates.

By Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang, Ziyang Ma, Ruiyang Xu, Meng Gao, Yinsong Yan, Ling Wang, Hui Wang, Wen Huang, Yiheng Chen, Guanrou Yang, Qiuqiang Kong, Jin Xu, Xie Chen
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Hugging Face Trending Papers
Jul 6

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.

arXiv AI
Sep 18

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.

By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
arXiv Computation and Language
Sep 11

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction

The paper introduces CoCoEval, a framework for evaluating large language model (LLM)–simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level. Using CoCoEval, the authors compare human conversations with those generated by GPT‑4.1, GPT‑5.1, and Claude Opus 4, finding that LLMs produce far fewer such behaviors under vanilla prompting and that prompt engineering or fine‑tuning often over‑produces specific behaviors. The study highlights gaps between human and LLM‑simulated interactions that conventional Likert‑scale evaluations miss, raising concerns about using LLMs as proxies for human social interaction.

By Ryo Kamoi, Ameya Godbole, Binglin Zhou, Xiaoxin Lu, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou
arXiv AI
2d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza