Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,971 stories · RSS feed

Hugging Face Trending Papers
Sep 28

One Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMU

The study investigates how well a single consumer earbud IMU can estimate 3D body pose and whether adding foot IMUs improves accuracy. Using a multimodal capture pipeline with RGB‑D video, an AirPods head IMU, and Striv insole IMUs, the authors benchmark pose estimation across various motions and train recurrent models (IMUPoser and MobilePoser). Results show that a head IMU alone achieves 79.0 mm rigid‑MPJPE and 0.809 macro‑F1 for foot contact, while adding foot IMUs does not significantly improve pose and can even degrade performance due to insole orientation errors.

Hugging Face Trending Papers
Sep 28

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

The paper introduces a training‑free visual token pruning strategy for vision‑language models that separates early vision‑guided pruning from later text‑guided reselection. By first pruning tokens with vision‑encoder attention, retaining candidates until the decoder midpoint, and then applying text‑to‑visual attention, the method preserves task‑relevant visual information. Across eight benchmarks and three models, it achieves an average performance recovery of 11.10 and 16.84 percentage points at 80% and 90% pruning, respectively, while maintaining comparable or lower LLM‑prefill latency.

arXiv Computer Vision
Sep 28

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

CCRV-Bench is a constraint‑driven benchmark designed to evaluate visual causal reasoning in vision‑language models on single‑image physical scenarios. It assesses four causal task dimensions—causal relation discovery, state prediction, causal diagnosis, and intervention—while applying constraints such as entity symbolization, spatial grounding, factual adversarial constraints, and minimalist output constraints to reduce shortcut learning. Experiments on 15 multimodal models reveal that constraint sensitivity varies by task and model, with intervention and spatial grounding having the largest impact and factual adversarial constraints improving causal diagnosis across models.

By Linyuan Gao, Yuan Wu, Yi Chang
arXiv Computer Vision
Sep 28

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

Auteur is a language‑driven method that generates human‑centric camera framing for generative video models. It treats shots as framings relative to an actor, encoding shot size, angle, and composition as functions of human pose and motion, and uses a domain‑specific language that converts to standard 6‑DoF camera parameters. A fine‑tuned multimodal large language model acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are interpolated into continuous camera trajectories for video generation.

By Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra, Xuelin Chen, Erkut Erdem, Aykut Erdem, Duygu Ceylan
arXiv Computer Vision
Sep 28

SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

SatNav is a new, scalable benchmark for long‑horizon vision‑language navigation (VLN) with unmanned aerial vehicles (UAVs), built from high‑resolution satellite imagery. It generates 118,000 navigation episodes across 59 scenes in 18 cities, using satellite crops to approximate UAV nadir views and featuring three task families—Boundary, Landmark, and Route—to test long‑term memory and geospatial reasoning. The benchmark also introduces SwiftVLN, a modular framework for memory component experimentation, and demonstrates that models trained on satellite data can transfer to real‑flight UAV observations.

By Jiajun Jiang, Chunliang Hua, Zichun Chen, Yanxing Wu, Zeyuan Yang, Jie Song, Xiao Hu
arXiv Computer Vision
Sep 28

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.

By Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Shichao Kan
arXiv AI
Sep 28

Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.

By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv AI
Sep 28

Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings

The paper introduces the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which fuses fMRI and MEG data to decode perceived speech from non‑invasive brain signals. Comprehensive experiments show that SICMD improves Top‑1, Top‑10, and Rankacc scores by over 10%, 10%, and 1.7% respectively, while cutting training costs by 88.8% and 60.5% compared to existing multi‑subject and intra‑subject approaches. Visualizations further confirm the method’s effectiveness.

By Aoke Zhang, Jing Chen
arXiv AI
Sep 28

Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting

The paper proposes Metric-based Loss Weighting to enhance visual grounding in multimodal machine translation. By increasing loss for tokens that benefit from image context—identified via the Point-wise Cross-mutual Information (PCXMI) metric and its Congruency-based variant—the method improves translation accuracy on the CoMMuTE dataset by over 7 percentage points. Experiments fine-tune three pretrained multimodal LLMs across three language directions, showing superior performance compared to standard fine-tuning while preserving overall translation quality.

By Pawe{\l} M\k{a}ka, Piotr Andruszkiewicz, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
arXiv AI
Sep 28

ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs

ViSTA is a lightweight adapter that adds irregular numerical measurements to a pretrained vision‑language model’s chart representations, enabling accurate clinical time‑series prediction without altering the base model’s parameters. On the MIMIC‑IV dataset, ViSTA outperforms other adaptations across four metrics for acute kidney injury and mortality prediction, achieving an AUC of 0.7376 with only 0.516 million trainable parameters. It also delivers strong temporal question‑answering performance, reaching 69.27% accuracy with significantly fewer trainable parameters than low‑rank adaptation methods.

By Junyi Gao, Yu Shi, Pingzhao Hu, Ewen M Harrison
arXiv AI
Sep 28

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

The paper introduces CaC, a coarse‑to‑fine anomaly reward model that uses Vision‑Language Models to first scan globally for anomalous time windows, then ground anomalies spatially, and finally reason with structured spatiotemporal Chain‑of‑Thought. It builds the first large‑scale generated video anomaly dataset with detailed annotations and trains the model through a three‑stage progressive paradigm, including reinforcement learning with Group Relative Policy Optimization. Experiments show CaC improves fine‑grained anomaly detection by 25.7% and reduces generated‑video anomalies by 11.7% while enhancing overall video quality.

By Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan, Boheng Zhang, Haonan Fan, Honglie Wang, Yiyang Fan, Zhenlong Yuan, Zijun Li, Yongrui Heng, Guosheng Lin, Fan Yang
arXiv AI
Sep 28

Prompt-Based Continual Compositional Zero-Shot Learning

The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.

By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv Computer Vision
Sep 28

Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment

The paper introduces ReasonCVS, a structured reasoning framework for assessing the Critical View of Safety in laparoscopic cholecystectomy. It uses a Vision‑Language Model to build an Anatomical Scene Graph Abstraction and a Large Language Model–based Rationale‑Aware Reasoning Agent to verify sub‑criteria, producing a final verdict with traceable clinical rationale. Experiments on the Endoscapes‑CVS201 benchmark show ReasonCVS outperforms existing methods with a 68.1% mAP while offering interpretable, criterion‑level explanations.

By Qing Xu, Yuxiang Luo, Zhen Chen
arXiv Computer Vision
Sep 28

InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

InternW0-Δ is a unified World Action Model that integrates pretrained visual dynamics, scene semantics, 4D geometry, and motion priors within a Mixture-of-Transformers framework to generate robot actions. It leverages a frozen VLM for semantic guidance, a 4D foundation model for geometric priors, and introduces Causal Imprint to learn future-relevant scene changes without future-video rollout. The model is pretrained on a newly curated 20K‑hour heterogeneous corpus of robot and human demonstrations, achieving superior performance on simulation benchmarks and real‑robot platforms.

By Xingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang, Hanxue Zhang, Yating Wang, Xudong Li, Yuping He, Xueyuan Wei, Chao Gao, Xijie Yang, Yingxiang Xu, Kerui Ren, Wenqi Guo, Jianjun Zhou, Xinzhe Wang, Weiguang Zhao, Ni Yang, Zetao Cai, Yufei Xue, Hengjie Li, Zeyu He, Yuanzhen Zhou, Rong Fu, Jianyang Zhang, Siwei Cui, Fuxian Huang, Yunsong Zhou, Xing Gao, Yifei Yao, Qiaojun Yu, Kailin Li, Ming Zhou, Mu Huang, Xinyue Li, Wenze Cui, Bingqi Jiang, Xueyue Zhu, Junting Dong, Haoyu Guo, Tao Lu, Mulin Yu, Bowen Zhou, Bin Zhao, Tianfan Xue, Weinan Zhang, Chunhua Shen
arXiv AI
Sep 28

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza
arXiv AI
Sep 28

What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

The paper investigates how different multimodal design choices affect the performance of misinformation detection systems. Using over 3,375 experiments across three benchmark datasets and various pre‑trained vision and language models, the authors systematically compare design options and conduct robustness analyses. The study offers practical guidance on which choices improve detection, when they may fail silently, and which pipeline components most influence model behavior, addressing four key research questions.

By Akshit Sharma, Prashant W. Patil
arXiv Computer Vision
Sep 28

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.

By Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
arXiv Computation and Language
Sep 28

Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition

The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.

By Bo Su, Yueru Yan, Thai Le
arXiv Computation and Language
Sep 28

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

I-Parakeet is an integer‑only implementation of NVIDIA’s Parakeet‑CTC Conformer ASR model that runs entirely on a smartphone NPU without any floating‑point operations or CPU fallback. The paper introduces three key techniques: an integer formulation of relative‑positional self‑attention, a minimax‑optimized Swish approximation, and layer‑wise range analysis with INT16 BatchNorm and percentile calibration for pre‑encoder activations. The resulting model achieves 4.97% WER on LibriSpeech test‑other and runs 7.5× faster than a CPU baseline on a Qualcomm NPU.

By Taichi Nishimura