arXiv:2609. 27303v1 Announce Type: new Abstract: Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself.
By Shujian Gao, Jiamei Yan, Yuchen Yang, Penghao Zhou, Qinglei Wang, Tiehan Fan, Yuan Wang, Zuxuan Wu, Yu-gang Jiang
The paper presents a benchmark to test whether vision‑language models can produce plant simulation configurations from images using in‑context learning. It focuses on cowpea plot reconstruction, requiring the models to output structured JSON that includes field and plant details. Open‑source multimodal models from the Gemma 4 and Qwen3.5 families are evaluated on synthetic and real drone datasets, using five in‑context methods, and the results show that while VLMs can generate valid JSON and estimate key agronomic metrics, their performance varies and often lags behind dataset baselines.
By Heesup Yun, Isaac Kazuo Uyehara, Earl Ranario, Lars Lundqvist, Christine H. Diepenbrock, Brian N. Bailey, J. Mason Earles
arXiv:2609.26867v1 Announce Type: cross
Abstract: This article is motivated by an imaging application from the Adolescent Brain Cognitive Development (ABCD) study, aiming to predict task-based brain...
By Rajarshi Guhaniyogi, Pritam Dey, Krishnendu Chandra, Aaron Scheffler, Bani K. Mallick
PR‑Smoother is an amortized smoothing method that preserves the explicit use of a prescribed simulator in both the evidence lower bound and the variational family. It learns only future‑conditioned corrections to the simulator’s rollout, yielding a non‑Gaussian smoothing distribution that can jointly infer state, parameters, and sensor bias from observations alone. The approach recovers the exact smoother in deterministic and linear‑Gaussian limits and has been shown to capture multimodal posteriors in Lorenz‑96 and scale to 16,384‑dimensional Kolmogorov flow.
By Yuta Tarumi
The paper proposes treating a neural network’s layers as time steps in a state‑space model, converting Bayesian training into a smoothing problem. By propagating Gaussian moments forward and applying a Rauch–Tung–Striebel backward pass, weight posteriors are updated in closed form without gradient iterations or replay. The authors extend prior work by introducing a cross‑covariance identity that allows full‑covariance propagation through nonlinear activations, enabling more accurate online adaptation in non‑stationary classification, dynamics learning, and vision‑language‑action policy adaptation.
By Oren Wright, Haoming Jing, Qiaoan Shen, Koichiro Niinuma, Yorie Nakahira, Jos\'e M. F. Moura
The paper investigates the Platonic Representation Hypothesis, which posits that more capable models converge toward shared representations. By distinguishing relational structure (which samples are related) from metric geometry (quantitative relations like distances), the authors develop a controlled $2 imes2$ framework to evaluate both aspects at local and global scales. Their findings show that relational structure consistently converges across vision‑language and video‑text models, while metric geometry converges much more weakly, a pattern that persists even when using a Riemannian metric approximation.
By Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung
VCMM: Variance-Calibrated Momentum for Multimodal Learning proposes a new optimizer that adapts momentum based on modality-specific gradient dynamics. It estimates minibatch noise and temporal drift online, using a Kalman-inspired controller to set modality-specific momentum and applies bias correction for the first moment. Experiments on four multimodal benchmarks show consistent improvements with modest training overhead.
By Zhongjing Gu, Chenyang Huang, Yufa Feng, Chong He, Qinxu Ding, Yiming Cui
The paper presents a dual‑matrix computational framework that quantifies morphological and theological differences among 196 Hindu and Vajrayana Buddhist esoteric deities. It uses a Gower distance matrix with a new Cardinality Weighting algorithm for physical form and dense vector embeddings from LLMs for theological function, revealing how visual forms can mask shared functions and how high‑cardinality symbols cluster orthodox and Tantric entities. The study demonstrates near‑identical coordinates for the Hindu Chinnamasta and Buddhist Chinnamunda, and releases the architecture as an open‑source tool for Digital Humanities research.
By Ankit Bhattacharjee
The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.
By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
arXiv:2609.27222v1 Announce Type: cross
Abstract: Pneumonia is difficult to diagnose in older long-term care residents; multimorbidity and atypical presentations obscure signs, motivating operational...
By Nicholas Rasmussen, Oleg Zaslavsky, Zih-Ling Wang, Hongyu Yu, Joelle Fathi, Kaibao Nie, Amil Khanzada, Tomoko Ito
arXiv:2609.26986v1 Announce Type: new
Abstract: For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference wi...
By Zhuoyun Li, Boxuan Wang, Xiaowei Huang, Yi Dong
arXiv:2609.27450v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale er...
By Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao
arXiv:2609.27816v1 Announce Type: cross
Abstract: Safe coordination in heterogeneous machine-to-machine (M2M) robotic systems is challenging when robots differ in sensing capabilities, environmental...
By Mohamed Dwedar, Ahmad Hafez, Alexander Jesser, Amr Alanwar
arXiv:2609.28049v1 Announce Type: cross
Abstract: Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as ev...
By Sai Varun Kodathala, Prashanth Pollishetty, Jaylen Cargill
arXiv:2609.28086v1 Announce Type: cross
Abstract: We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings...
By Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur
arXiv:2609.27195v1 Announce Type: cross
Abstract: Non-speech interference can change a speech representation without causing comparable task loss. We test eight frozen encoders on four tasks, adding...
By Vsevolod Kovalev, Pranay Manocha
arXiv:2609.26907v1 Announce Type: cross
Abstract: Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal cla...
By Akshit Sharma, Prashant W. Patil
arXiv:2609.27848v1 Announce Type: new
Abstract: The BIC-MAC challenge targets whole-body pseudo-CT synthesis from NAC-PET, Dixon MRI, and a 2D topogram for CT-less PET attenuation correction. We prop...
By Xuan Loc Nguyen, Hoang-Loc Cao, Truong Thanh Hung Nguyen, Phuc Ho, Phuc Truong Loc Nguyen, Nguyen Truong Toan To, Hung Cao
arXiv:2609.28222v1 Announce Type: new
Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
By Xueqi Qiu, Xingyu Miao, Jingjing Deng, Haoran Duan, Yang Long, Ling Shao
arXiv:2609.28236v1 Announce Type: new
Abstract: Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounte...
By Lizhou Liang, Xinyu Zhong, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Qinfeng Li, Peng Li, Jintao Chen, Xuhong Zhang, Wenqi Zhang