arXiv:2609.24576v1 Announce Type: cross
Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...
By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv:2509.22650v3 Announce Type: replace
Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
By Anna Kukleva, Enis Simsar, Alessio Tonioni, Muhammad Ferjad Naeem, Federico Tombari, Jan Eric Lenssen, Bernt Schiele
arXiv:2606.07338v2 Announce Type: replace
Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rational...
By Zikai Zhang, Hubert P. H. Shum, Toby P. Breckon
arXiv:2608.10706v3 Announce Type: replace
Abstract: Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surf...
By Shuai Wang, Wangyuan Ding, Yixian Shen, Jia-Hong Huang, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring
arXiv:2609.07780v2 Announce Type: replace
Abstract: Automated drone surveillance has become increasingly important for public safety, critical infrastructure protection,and restricted airspace monito...
By Ami Pandat, Rajasekhar Punna, Gopika Vinod, Rohit Shukla
M3GA-Wild is a new benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests, combining synchronized RGB imagery and LiDAR from ground traversals with high‑resolution aerial imagery and multi‑altitude LiDAR over 370 hectares. The dataset includes accurate geo‑referenced 6‑DoF poses and spans 36 km of forest traversals, enabling systematic evaluation of visual, LiDAR, cross‑modal, and multi‑modal methods. Baseline experiments show LiDAR outperforms vision‑only approaches under severe viewpoint changes, while current multi‑modal fusion offers limited gains due to poor cross‑modal alignment, highlighting challenges in cross‑platform localisation and domain gaps.
By Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani
BrainIAC is a unified framework for 3D brain lesion segmentation that handles heterogeneous MRI modalities and adapts online during interactive segmentation. It combines a multi‑modal backbone trained with zero‑filling and random modality dropping, 3D interactive prompts that default to fully automatic predictions, and a two‑stage online adaptation guided by pseudo‑labels and a Click‑Centered Gaussian loss. Experiments on seven MRI datasets show that the components work synergistically, outperforming existing methods and generalizing to unseen modalities and pathologies.
By Wentian Xu, Anthony P Addison, Ziyun Liang, Harry Anthony, Guang Yang, Konstantinos Kamnitsas
The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.
By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
DeCo introduces an efficient decouple-to-couple learning framework for multi-task visual grounding, addressing conflicts between localization and segmentation tasks. It first applies Task-aware Semantic Decoupling (TSD) to separate shared visual cues into task-specific features guided by salient words, then uses Hybrid Prior Coupling (HPC) to merge sentence-level semantic priors with mask-derived spatial priors for improved grounding. Experiments across multiple natural and remote sensing datasets show that DeCo achieves state‑of‑the‑art performance while requiring only lightweight trainable parameters on a frozen multimodal encoder.
By Xiaoqiang Lu, Licheng Jiao, Long Sun, Yuting Yang, Xu Liu, Lingling Li, Wenping Ma, Fang Liu
SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.
By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang
CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.
By Gang Liu, Yanling Hao, Yixuan Zou
The CXR‑LT 2026 Challenge introduces a multi‑center, long‑tailed chest X‑ray classification benchmark with over 145,000 radiologist‑annotated images from PadChest and NIH datasets. It defines two core tasks: robust multi‑label classification on 30 known classes and open‑world generalization to 6 unseen rare disease classes. The paper outlines data collection, annotation, solution strategies, and evaluates performance across head‑vs‑tail, calibration, and cross‑center gaps, noting that vision‑language models improve in‑distribution and zero‑shot performance but rare‑finding detection under multi‑center shift remains difficult.
By Hexin Dong, Yi Lin, Pengyu Zhou, Fengnian Zhao, Alan Clint Legasto, Juno Cho, Dohui Kim, Justin Namuk Kim, Mingeon Kim, Sunwoo Kwak, Gabriel Moy\`a-Alcover, Ky Trung Nguyen, Thanh-Huy Nguyen, Ha-Hieu Pham, Huy-Hieu Pham, Huy Le Pham, Nikhileswara Rao Sulake, Aina Tur-Serrano, Ruichi Zhang, Ang Zu, Adam E. Flanders, Zhiyong Lu, Ronald M. Summers, Mingquan Lin, Hao Chen, Yuzhe Yang, George Shih, Yifan Peng
MinCU is a new benchmark for grounded minimal‑change understanding that presents pairs of near‑identical images differing by a single atomic variation in object category, attribute, count, or spatial position. Models are evaluated on their ability to describe the change, localize the changed region, and identify the changed entity. The authors also introduce SG‑ISA, a structured autoregressive method that decomposes the task into a Think‑Locate‑Describe sequence, showing that fine‑tuning with SG‑ISA improves both grounding accuracy and description quality while reducing reasoning‑token overhead.
By Chaoqian Mu, Wenhao Wu, Zichen Liang, Jiaxu Li, Lijun Wang, Yifan Wang, Huchuan Lu
The study developed a multimodal ultrasound and clinical data model to predict microvascular invasion (MVI) preoperatively in hepatocellular carcinoma (HCC). Using data from 489 patients across eight centers, the model combined B-mode ultrasound, color Doppler flow imaging, dynamic contrast-enhanced ultrasound, and clinical information, achieving an AUC of 0.8953 in external validation. Dynamic contrast-enhanced ultrasound contributed the most predictive power, while other modalities and clinical data added complementary value.
By Jun Cheng, Yuanyuan Kong, Qing Huang, Xiaotong Tan, Licong Dong, Yulong Han, Wufeng Xue, Ruobing Huang, Dong Ni, Qi Yang, Jie Yu, Ping Liang
LIBERO-VPro is a benchmark designed to assess the closed‑loop visual robustness of robotic foundation models by systematically perturbing visual inputs during task execution. It spans four dimensions—Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task‑Relevant Scene Variation—across 12 challenge categories, 96 settings, and 3,296 task‑condition cases. Evaluations on six models over 196,000 simulated episodes and 200 real‑world rollouts show that high nominal performance can hide significant weaknesses in visual grounding, adaptation, and sensitivity to stale or missing observations.
By Huiqiong Li, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Yu-Gang Jiang, Bin Zhu
The study reports that agitation—an early, subtle sign of distress—can be detected in autistic youth using multimodal wearable sensors. By fusing upper‑body movement, wrist‑worn physiological data, and vocalizations through pretrained foundation models, the authors achieved an AUC of 0.724 at the clinician‑annotated onset of agitation, with audio contributing the most signal. The approach works across 15 participants, showing that individualized agitation is detectable even before the onset and without needing a separate model per child.
By Nibraas Khan, Abigale Plunk, John Staubitz, Ingrid Shragge, Jordan Brooks, Suzanne Wright, Alec Brewer, James Dieffenderfer, Alper Bozkurt, Amy Weitlauf, Nilanjan Sarkar
arXiv:2609.24579v1 Announce Type: new
Abstract: Event logs arise in a wide range of real-world processes, capturing not only event activities and timestamps but also multi-modal contextual informatio...
By Fabian Spaeh, Jingxing Fang, Shandian Zhe, Bin Shen
arXiv:2609.24718v1 Announce Type: new
Abstract: While a centralized approach involving patient consent to collect and analyze data centrally would theoretically offer the best data quality and predic...
By Anne-Christin Hauschild, Amirreza Aleyasin, Nils H. Beyer, Lisa Fricke, Jonas H\"ugel, Maryam Moradpour, Anh-Tien Nguyen, Youngjun Park, Sophia Rheinl\"ander, Tim Beissbarth, Elisabeth Hessmann, Martin Middeke, Matthias Lauth, Maximilian Reichert, Ulrich Sax
arXiv:2609.22094v1 Announce Type: cross
Abstract: Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining fo...
By Zeeshan Ahmed, Yang Qin, Hanqing Huang
arXiv:2609.22149v1 Announce Type: cross
Abstract: Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after ta...
By Yefeng Yuan, Zhan Shi, Liang Cheng, Yuhong Liu