Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are enta...
Microsoft Research introduces Quine, an early‑stage AI research system that builds a multimodal world model of biology. By linking insights across different biological scales and modalities, Quine enables scientists to computationally explore a much larger hypothesis space than intuition alone would allow. Experimental results feed back into the system, helping researchers refine future research directions.
By Nicolo Fusi, Jonathan M. Carlson
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teac...
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, re...
Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model...
While Large Vision-Language Models (LVLMs) achieve remarkable success, hallucinations remain a significant barrier to their reliable deployment. Recent studies primarily attribute these issues to cros...
Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general o...
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover de...
Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in...
Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.
By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
LUMO (Lightweight Unified Multilingual Orchestrator) is a privacy‑preserving offline voice assistant that runs entirely on edge hardware, specifically a Raspberry Pi 5 with 8 GB RAM. It integrates local ASR, a 4‑bit GGUF‑quantized LLM, and TTS to deliver end‑to‑end response latencies of 2.0–4.0 s, a 6.8 % WER on short English utterances, and lower peak power consumption (~9 W) compared to existing edge assistants. The system also supports Bangla speech, enabling multilingual use in low‑resource settings.
By Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
WorldTS is a new forecasting framework that models latent dynamics conditioned on multimodal covariates to improve time‑series prediction. It uses a two‑stage training process: first learning latent state dynamics from historical data and covariates, then training a decoder to map predicted latent states back to future observations. Experiments on 21 real‑world datasets demonstrate the effectiveness of this approach.
By Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen
The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.
By Diego Torres, Axel Roebel, Nicolas Obin
The paper introduces the Neural Action Codec (NAC), a convolutional encoder‑decoder architecture that treats short robot action trajectories as multi‑channel 1D signals and compresses them using a multi‑scale residual vector quantization (RVQGAN) model. NAC replaces traditional discrete action tokenizers with a compact, ordered token space via offset codebooks, allowing standard autoregressive policies to operate over short, structured sequences while a Vocos‑style decoder reconstructs the actions. Experiments on LIBERO‑10, RoboMimic, and real‑world manipulation tasks show that NAC achieves higher reconstruction fidelity and better average success rates than existing binning, FAST, and VQ‑based tokenizers at comparable or improved compression rates.
By Ahad Jawaid, Yu Xiang
EAServe introduces an encode-aware disaggregated serving framework for multimodal large language models (MLLMs), restructuring the traditional Prefill-Decode pipeline into a three-stage Encode-Prefill-Decode (EPD) system. By treating Encode as the control point, EAServe coordinates load‑adaptive micro‑batching, rate‑controlled offloading to prefill workers, and dynamic SM partitioning to balance GPU utilization across stages. Its Hybrid Auto Selection (HAS) layer optimizes GPU allocation, encode batch size, and offload ratio using capacity profiling and Bayesian optimization, achieving up to 4.3× higher goodput compared to NVIDIA Dynamo and 1.7× higher than vLLM on various MLLM architectures.
By Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin, Gagan Agrawal, Wei Niu
ChemMLLM is a unified chemical multimodal large language model designed for molecule understanding and generation across text, SMILES strings, and images. The authors curated five multimodal tasks and benchmarked ChemMLLM against leading general MLLMs, chemical LLMs, and specialized models, finding it outperforms general-purpose MLLMs and matches specialized models on all tasks. The study demonstrates that a single foundation model can handle diverse cross‑modal chemical tasks, including image generation, enabling more intuitive visual human‑AI interaction.
By Qian Tan, Di Zhang, Ben Gao, Peng Xia, Wanhao Liu, Shufei Zhang, Wanli Ouyang, Lei Bai, Yuqiang Li, Tianfan Fu
The paper introduces EmoSpeechBrain, a multimodal emotion recognition framework that fuses EEG and speech signals. It employs differential attention in the EEG encoder to cancel shared noise and an attention-based gating adapter to align modalities and weight their contributions. Experiments on PME4 and EAV datasets show up to 12.9% accuracy improvement over other EEG encoders and surpass unimodal baselines by up to 23.1%.
By Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
The paper introduces Supervised Deep Multimodal Matrix Factorization (SD3MF), an interpretable framework that extends Symmetric Nonnegative Matrix Tri-Factorization to supervised prediction across populations of multimodal brain graphs. SD3MF learns deep hierarchical factorizations for each modality and a shared latent representation, jointly optimizing graph reconstruction and prediction while enabling data-driven multimodal fusion. Experiments on multimodal connectome datasets demonstrate that SD3MF outperforms strong deep learning baselines such as CNNs and GNNs, providing biologically interpretable insights through community-level interaction matrices.
By Amjad Seyedi, Lifang He, Songlin Zhao, Akwum Onwunta, Nicolas Gillis
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support t...
BiMoGen introduces a unified masked discrete diffusion framework for bidirectional motion‑text generation, addressing the limitations of autoregressive models in capturing bidirectional dependencies between language and motion. The approach employs a two‑stage training strategy—decoupled uni‑ and cross‑modal pretraining followed by supervised fine‑tuning—to establish robust cross‑modal correspondence, and incorporates Generation‑Aware Self‑Correction to mitigate error propagation during inference. Experiments on HumanML3D and KIT‑ML show competitive performance on both text‑to‑motion and motion‑to‑text tasks, demonstrating the effectiveness of the proposed training and correction mechanisms.