arXiv AI

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

arXiv:2509. 09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline.

arXiv AI
Jun 2

Multimodal Music Recommendation System using LLMs

arXiv:2606. 00125v1 Announce Type: cross Abstract: Music recommendation systems typically treat songs as opaque tokens, relying on collaborative interaction histories which overlooks semantic or acoustic content.

By Srikar Prabhas Kandagatla, Sreehitha R. Narayana, Chandana Magapu, Swetha Mohan, Shamanth Kuthpadi, Hongjie Chen, Ryan A. Rossi, Franck Dernoncourt, Nesreen Ahmed
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv AI
Aug 24

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

The paper surveys how large multimodal models (LMMs) enhance agentic frameworks that combine perception, memory, reasoning, planning, and action. It examines the integration of multiple modalities—text, images, audio, and video—through delegated, late‑fusion, and early‑fusion architectures, and maps these designs to agent capabilities. The survey also reviews multimodal agentic systems in robotics, web navigation, multimedia content creation, and video understanding, evaluating performance, efficiency, and scalability trade‑offs.

By Neel Mokaria, Rishie Raj, Dheeraj Baiju, Xiaoqian Shen, Shraman Pramanick, Kevin Qinghong Lin, Arda Senocak, Mike Zheng Shou, Philip Torr, Mohamed Elhoseiny, Yapeng Tian, Ruohan Gao, Salman Khan, Sayan Nag, Sanjoy Chowdhury, Dinesh Manocha
arXiv AI
6d ago

Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops

The paper presents a pipeline for generating multi‑turn synthetic conversations and a self‑improvement loop that uses variance‑based contrastive optimization and a coding agent to refine planning and tool‑use in conversational recommendation agents. This approach improves agent quality by 8% over a manually optimized prompt and has been deployed at Spotify, where it accelerated development cycles. In production, the system achieved a 14% increase in user listening, a 5% rise in weekly active users, and a 5% reduction in skip rate compared to a prior session‑only experience.

By Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galv\~ao, Hugues Bouchard, Mounia Lalmas, Jos\'e Luis Redondo Garc\'ia, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindstr\"om, Dani Doro, Christine Doig Cardet
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
arXiv AI
Sep 21

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

MENASpeechBank is a new reference voice bank that provides about 18,000 high‑quality utterances from 124 speakers across multiple MENA countries, covering English, Modern Standard Arabic, and regional Arabic varieties. The dataset is built through a controllable pipeline that creates persona profiles inspired by the World Values Survey, defines a taxonomy of roughly 5,000 conversational scenarios, matches personas to scenarios via semantic similarity, and generates around 417,000 role‑play conversations using an LLM. Synthetic speaker‑conditioned user‑turn audio is produced from reference recordings to maintain speaker diversity, and both synthetic and human‑recorded conversations are evaluated and analyzed for quality.

By Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam
arXiv AI
2d ago

MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

arXiv:2505.17613v2 Announce Type: replace Abstract: Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation...

By Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, Ranjay Krishna, Lucy Lu Wang, Yulia Tsvetkov, Noah A. Smith, Banghua Zhu
arXiv AI
2d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
arXiv AI
Jun 12

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.

By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos