arXiv:2609.27094v1 Announce Type: cross
Abstract: Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inheren...
By Alexandros Alexiou, Charilaos Papaioannou, Alexandros Potamianos
arXiv:2605.16080v2 Announce Type: replace
Abstract: The rise of AI-generated images (AIGIs) poses growing challenges for digital authenticity, prompting the need for efficient, generalizable image fo...
By Qing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu, Jian Zhang
arXiv:2609.23407v2 Announce Type: replace-cross
Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challengin...
By Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong
arXiv:2508.13680v5 Announce Type: replace-cross
Abstract: We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (V...
By Vy Tuong Dang, An Vo, Emilio Villa-Cueva, Quang Tau, Duc Dm, Thamar Solorio, Daeyoung Kim
The article presents design principles and a software architecture, AGIMUD, for enabling interaction between humans and multiple agents in dynamic simulated worlds. It integrates socially-aware reasoning, emotional agent behavior, a multimodal human interface, and distributed AI processing to support real‑time, multi‑user dungeon (MUD) environments. The work builds on current AI/AGI and transformer‑based conversational agents to create sustainable, governance‑aware human‑agent reasoning systems.
By David Berga
Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmenta...
Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language fe...
Google has launched two new Gemini text‑to‑speech models—gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts—offering a library of over 2,000 voices and the option to create a custom voice from a 30‑second audio sample. The author built a playground interface that lets users define multi‑character conversations with distinct voices and styles, and demonstrated it with a scripted dialogue between two pelicans. Generating 1 minute 18 seconds of audio with the Flash model took about 20 seconds and cost 2.74 cents.
Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of...
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Ama...
Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the...
Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the sa...
Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{co...
In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with spea...
Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,...
Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a...
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit g...
Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language mode...
arXiv:2609.24124v1 Announce Type: cross
Abstract: Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effect...
By Yibo Li, Enshen Zhou, Rui Chen, Yanjun Ding, Mengzhen Liu, Yi Han, Jiabo Zhan, Lipeng Wang, Shanghang Zhang, Lu Sheng