Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,022 stories · RSS feed

arXiv AI
Jul 23

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

arXiv:2607. 20345v1 Announce Type: cross Abstract: Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability.

By Roger Sala Sis\'o, Tiago Silv\'erio, Jakob Sand, Tran Nguyen Le
arXiv Machine Learning
Jul 23

Missing-by-Design: Certifiable Modality Deletion for Revocable Multimodal Sentiment Analysis

arXiv:2602. 16144v4 Announce Type: replace-cross Abstract: As multimodal systems increasingly process sensitive personal data, the ability to selectively revoke specific data modalities has become a critical requirement for privacy compliance and user autonomy.

By Rong Fu, Ziming Wang, Chunlei Meng, Jiekai Wu, Kangan Qian, Hao Zhang, Simon Fong
arXiv AI
Jul 23

SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection

arXiv:2604. 26633v2 Announce Type: replace-cross Abstract: Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, and collecting balanced training sets is slow and costly.

By Paul Julius K\"uhn, Mika Pommeranz, Arjan Kuijper, Saptarshi Neil Sinha
arXiv AI
Jul 23

SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework

arXiv:2602. 17330v5 Announce Type: replace-cross Abstract: Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost of pairwise affinity evaluations and dataset imbalances that obscure clinically important minority clonotypes.

By Rong Fu, Zijian Zhang, Kun Liu, Jiekai Wu, Xianda Li, Simon Fong
arXiv AI
Jul 23

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection

arXiv:2512. 16300v3 Announce Type: replace Abstract: Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language models (MLLMs) with high-level semantic knowledge.

By Fanrui Zhang, Qiang Zhang, Sizhuo Zhou, Jianwen Sun, Chuanhao Li, Jiaxin Ai, Yukang Feng, Yujie Zhang, Wenjie Li, Zizhen Li, Yifan Chang, Jiawei Liu, Kaipeng Zhang
arXiv Machine Learning
Jul 23

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

arXiv:2607. 19816v1 Announce Type: cross Abstract: Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules.

By Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li, Boxuan Zhao, Yonghong Tian, Bartosz A. Grzybowski, Fanyang Mo