Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv Machine Learning
Sep 24

Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams

arXiv:2609. 27303v1 Announce Type: new Abstract: Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself.

By Shujian Gao, Jiamei Yan, Yuchen Yang, Penghao Zhou, Qinglei Wang, Tiehan Fan, Yuan Wang, Zuxuan Wu, Yu-gang Jiang
arXiv AI
Sep 24

Using Vision Language Foundation Models to Generate Plant Simulation Configurations via In-Context Learning

The paper presents a benchmark to test whether vision‑language models can produce plant simulation configurations from images using in‑context learning. It focuses on cowpea plot reconstruction, requiring the models to output structured JSON that includes field and plant details. Open‑source multimodal models from the Gemma 4 and Qwen3.5 families are evaluated on synthetic and real drone datasets, using five in‑context methods, and the results show that while VLMs can generate valid JSON and estimate key agronomic metrics, their performance varies and often lags behind dataset baselines.

By Heesup Yun, Isaac Kazuo Uyehara, Earl Ranario, Lars Lundqvist, Christine H. Diepenbrock, Brian N. Bailey, J. Mason Earles
arXiv Machine Learning
Sep 24

PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation

PR‑Smoother is an amortized smoothing method that preserves the explicit use of a prescribed simulator in both the evidence lower bound and the variational family. It learns only future‑conditioned corrections to the simulator’s rollout, yielding a non‑Gaussian smoothing distribution that can jointly infer state, parameters, and sensor bias from observations alone. The approach recovers the exact smoother in deterministic and linear‑Gaussian limits and has been shown to capture multimodal posteriors in Lorenz‑96 and scale to 16,384‑dimensional Kolmogorov flow.

By Yuta Tarumi
arXiv Machine Learning
Sep 24

Full-Covariance Smoothing of Bayesian Neural Networks for Online Adaptation

The paper proposes treating a neural network’s layers as time steps in a state‑space model, converting Bayesian training into a smoothing problem. By propagating Gaussian moments forward and applying a Rauch–Tung–Striebel backward pass, weight posteriors are updated in closed form without gradient iterations or replay. The authors extend prior work by introducing a cross‑covariance identity that allows full‑covariance propagation through nonlinear activations, enabling more accurate online adaptation in non‑stationary classification, dynamics learning, and vision‑language‑action policy adaptation.

By Oren Wright, Haoming Jing, Qiaoan Shen, Koichiro Niinuma, Yorie Nakahira, Jos\'e M. F. Moura
arXiv Machine Learning
Sep 24

What Converges in the Platonic Representation Hypothesis? Structure over Geometry

The paper investigates the Platonic Representation Hypothesis, which posits that more capable models converge toward shared representations. By distinguishing relational structure (which samples are related) from metric geometry (quantitative relations like distances), the authors develop a controlled $2 imes2$ framework to evaluate both aspects at local and global scales. Their findings show that relational structure consistently converges across vision‑language and video‑text models, while metric geometry converges much more weakly, a pattern that persists even when using a Riemannian metric approximation.

By Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung
arXiv Machine Learning
Sep 24

VCMM: Variance-Calibrated Momentum for Multimodal Learning

VCMM: Variance-Calibrated Momentum for Multimodal Learning proposes a new optimizer that adapts momentum based on modality-specific gradient dynamics. It estimates minibatch noise and temporal drift online, using a Kalman-inspired controller to set modality-specific momentum and applies bias correction for the first moment. Experiments on four multimodal benchmarks show consistent improvements with modest training overhead.

By Zhongjing Gu, Chenyang Huang, Yufa Feng, Chong He, Qinxu Ding, Yiming Cui
arXiv Machine Learning
Sep 24

Quantifying the Occult: A Comparative Study of Hindu and Buddhist Deities Using Machine Learning Methods

The paper presents a dual‑matrix computational framework that quantifies morphological and theological differences among 196 Hindu and Vajrayana Buddhist esoteric deities. It uses a Gower distance matrix with a new Cardinality Weighting algorithm for physical form and dense vector embeddings from LLMs for theological function, revealing how visual forms can mask shared functions and how high‑cardinality symbols cluster orthodox and Tantric entities. The study demonstrates near‑identical coordinates for the Hindu Chinnamasta and Buddhist Chinnamunda, and releases the architecture as an open‑source tool for Digital Humanities research.

By Ankit Bhattacharjee
arXiv Machine Learning
Sep 24

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.

By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
arXiv AI
Sep 24

BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

arXiv:2609.27450v1 Announce Type: cross Abstract: Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale er...

By Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao
arXiv AI
Sep 24

LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations

arXiv:2609.28086v1 Announce Type: cross Abstract: We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings...

By Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur
arXiv Computer Vision
Sep 24

Geometry-anchored PET-aware multimodal pseudo-CT synthesis for whole-body attenuation correction: the BIC-MAC Challenge

arXiv:2609.27848v1 Announce Type: new Abstract: The BIC-MAC challenge targets whole-body pseudo-CT synthesis from NAC-PET, Dixon MRI, and a 2D topogram for CT-less PET attenuation correction. We prop...

By Xuan Loc Nguyen, Hoang-Loc Cao, Truong Thanh Hung Nguyen, Phuc Ho, Phuc Truong Loc Nguyen, Nguyen Truong Toan To, Hung Cao