Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,974 stories · RSS feed

arXiv Computer Vision
Sep 23

MatchFusion: Explicit-Implicit Instance Matching for Spatio-Temporal Multimodal Autonomous Driving

MatchFusion is a learnable module that performs explicit-implicit instance matching for spatio‑temporal multimodal autonomous driving. It initializes pairwise affinities with geometric similarity and category consistency, then refines associations using instance embeddings to guide a residual aggregation operator for adaptive information exchange. Experiments on nuScenes show that MatchFusion improves perception accuracy, reduces FLOPs by 55.3% and GPU memory usage by 39.3%, and adds only 3.7% of total perception latency.

By Xiaoyu Li, Jiajia Fu, Long Shi, Tianyu Du, Ruihang Li, Xian Wu, Lijun Zhao, Yingtao Zhang, Lining Sun, Ruifeng Li
arXiv Computer Vision
Sep 23

GRIP: Gaussian Rendering as a Cross-Modal Bridge for Image-to-Point Cloud Registration

GRIP is a pose‑conditioned refinement framework that improves pixel‑to‑point matching for image‑to‑point‑cloud registration. It mitigates the mismatch between grid‑based image descriptors and unordered point cloud descriptors by softly rendering learned 3D point features onto the image grid using Gaussian feature splatting. The resulting rendered point‑derived feature map is fused with image features via a pixel‑aligned transformer, enabling visual semantic and geometric cues to interact in a shared 2D representation, which is then decoded and propagated to finer resolutions for dense correspondence estimation and final pose refinement. Experiments on RGB‑D Scenes V2 and 7 Scenes show state‑of‑the‑art inlier ratios and competitive registration recall, especially under stricter evaluation thresholds.

By Karim Slimani, Catherine Achard, Eric Marchand, Brahim Tamadazte
arXiv AI
Sep 23

HybridFlow: A 2-NFE Generative Policy for Real-Time Robotic Manipulation

HybridFlow is a generative policy for robotic manipulation that uses a three‑stage inference procedure requiring only two network function evaluations (2‑NFE). The policy first generates a coarse action trajectory with a Global Jump based on MeanFlow, then refines the state using a parameter‑free ReNoise interpolation, and finally performs a Local Refine to query the instantaneous‑velocity limit. Experiments on RoboMimic and five real‑robot settings show that HybridFlow achieves high success rates and improves task performance over a 16‑step Diffusion Policy while reducing action‑generation latency by roughly eightfold.

By Zhenchen Dong, Fulin Chen, Jinna Fu, Jiaming Wu, Qingran Wu, Shengyuan Yu, Hongyu Yu, Yide Liu
arXiv AI
Sep 23

SAIL: Test-Time Scaling for In-Context Imitation Learning with VLM

SAIL is a framework that transforms robot imitation learning into an iterative refinement problem, enabling test-time scaling of trajectory generation. It employs Monte Carlo Tree Search where each node represents a full trajectory and edges denote refinements, guided by an archive of successful trajectories, a vision‑language model for scoring, and step‑level feedback. Experiments on six manipulation tasks in simulation and real‑world settings show that higher test‑time compute consistently raises success rates, reaching up to 95% on complex tasks.

By Makoto Sato, Yusuke Iwasawa, Yujin Tang, So Kuroki
arXiv Machine Learning
Sep 23

LiveProBench: Can Streaming Video Models Really Interact Like Humans?

LiveProBench evaluates streaming video models on their ability to interact proactively, assessing whether they respond at appropriate times without explicit cues. The benchmark tests models at one‑second intervals across six subtasks that vary trigger ambiguity and timing tolerance, measuring response accuracy, silence rates, and duplicate responses. Results show that many models issue premature responses more often than missed ones, highlighting a significant shortfall in human‑like temporal decision making.

By Kaixuan Du, Xin Wan, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, YuKun Wang
arXiv Computation and Language
Sep 23

Same Chart, Different Story: Bias in Vision-Language Chart Interpretation

The paper introduces ChartBias, a benchmark of 820 real-world charts covering six social attributes, designed to audit bias in vision‑language models (VLMs) that interpret charts. Across 12 VLMs, the study identifies three failure modes—narrative shift, group hallucination, and preference polarity—where models produce different or misleading narratives when the referenced social group changes. A multi‑agent mitigation framework is proposed, separating evidence extraction from group‑conditioned generation and using a counterfactual judge, which reduces narrative shift while maintaining chart‑grounded reasoning.

By Mizanur Rahman, Huan Wu, Arash Asgari, Enamul Hoque Prince, Laleh Seyyed-Kalantari
arXiv Computation and Language
Sep 23

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI: A Narrative Review and Conceptual Framework

The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.

By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv Computer Vision
Sep 23

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Metric-Bench introduces a new benchmark for Vision‑Language Models (VLMs) that focuses on metric‑spatial reasoning in indoor scenes by using in‑image reference objects with known dimensions. The accompanying MetricReasoner fine‑tuning recipe employs structured prompts and numerical rewards to implicitly learn 2D‑to‑3D mapping without camera intrinsics. Experiments show that this approach improves spatial metric understanding by 43.1 % over existing models and boosts downstream embodied tasks, while also delivering gains on general VLM benchmarks.

By Yuling Xi, Haokai Zhang, Muzhi Zhu, Hao Zhong, Zongze Du, Hengyu Zhao, Chenchen Jing, Yufei Yin, Bin Qin, Yongjie Yang, Zhenbo Luo, Hao Chen, Chunhua Shen
arXiv AI
Sep 23

DiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic Simulation

arXiv:2609.23103v1 Announce Type: cross Abstract: While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess phys...

By Guanxiong Chen, Yiduo Qu, Qianjun Xia, Pengyu Jing, Yixian Cheng, Bole Ma, Pengzhi Yang, Bingyang Zhou, Ziming Li, Shashwat Suri, Gongbo Sun, Chao Liu, Peter Yichen Chen, Ziqiu Zeng, Fan Shi
arXiv Computer Vision
Sep 23

Moving6DPoSe: A Multimodal Database for Monocular 6D Pose Estimation and Segmentation of Moving Objects

Moving6DPoSe is a multimodal database for monocular 6D pose estimation and segmentation of moving objects, comprising two subsets: real-world recordings (Moving6DPoSe‑R) and synthetic sequences (Moving6DPoSe‑S). It includes 16 scanned objects, 1,702 real and synthetic rosbags, and annotations for semantic segmentation, object detection, and monocular 6D pose estimation. Baseline results show that event-based representations outperform conventional RGB images for moving‑object segmentation, while monocular orientation estimation remains challenging.

By Ignacio Bugueno-Cordova, Javier Ruiz-del-Solar, Rodrigo Verschae