Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv AI
Sep 24

OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents

arXiv:2609.23407v2 Announce Type: replace-cross Abstract: Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challengin...

By Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong
arXiv AI
Sep 24

Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations

The article presents design principles and a software architecture, AGIMUD, for enabling interaction between humans and multiple agents in dynamic simulated worlds. It integrates socially-aware reasoning, emotional agent behavior, a multimodal human interface, and distributed AI processing to support real‑time, multi‑user dungeon (MUD) environments. The work builds on current AI/AGI and transformer‑based conversational agents to create sustainable, governance‑aware human‑agent reasoning systems.

By David Berga
Simon Willison
Sep 23

Gemini 3.8 TTS Playground

Google has launched two new Gemini text‑to‑speech models—gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts—offering a library of over 2,000 voices and the option to create a custom voice from a 30‑second audio sample. The author built a playground interface that lets users define multi‑character conversations with distinct voices and styles, and demonstrated it with a scripted dialogue between two pelicans. Generating 1 minute 18 seconds of audio with the Flash model took about 20 seconds and cost 2.74 cents.