arXiv:2609.26097v1 Announce Type: cross
Abstract: Remote-sensing (RS) multimodal large language models (MLLMs) are trained and evaluated only in English, while text-only instruction data covers over...
By Xuechen Li
arXiv:2609.26658v1 Announce Type: cross
Abstract: Integrating heterogeneous datasets within data lakes is a critical challenge, particularly for semantically related tables that lack the explicit att...
By Md Ataur Rahman, Dimitris Sacharidis, Oscar Romero, Sergi Nadal
arXiv:2609.14711v2 Announce Type: replace
Abstract: Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport pro...
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv:2601.20900v4 Announce Type: replace-cross
Abstract: Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet...
By Andr\'es Carofilis, Sergio Burdisso, Esa\'u Villatoro-Tello, Shashi Kumar, Kadri Hacioglu, Srikanth Madikeri, Pradeep Rangappa, Manjunath K E, Petr Motlicek, Shankar Venkatesan, Andreas Stolcke
arXiv:2609.25270v1 Announce Type: new
Abstract: Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving b...
By Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
arXiv:2609.25615v1 Announce Type: new
Abstract: Architectural floor plans remain a high-friction barrier to archive digitization and early design-model preparation because heterogeneous graphics enco...
By Hongxuan Chen, Wenda Wang, Jiachen Lu, Qirui Shen, Zilong Huang, Lei He, Xinyue Dong, Weixin Huang
arXiv:2609.25850v1 Announce Type: new
Abstract: Deep learning performance generally improves with increasing training data, yet this scaling is fundamentally constrained by annotation cost in large-s...
By Xiaofei Du, Lei Zhang, Shuyu Yan, Manning Wang, Zhijian Song
arXiv:2609.25945v1 Announce Type: new
Abstract: The field of Artificial Intelligence has been adopted for many application domains. Vision Language Models are one of the recently advanced AI techniqu...
By Malsha Ashani Mahawatta Dona, Konstantinos Rokanas, Alexander S\"afstr\"om, Krishna Ronanki, Christian Berger
arXiv:2609.26093v1 Announce Type: new
Abstract: Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-g...
By Feixiang Liu, Qiang Qiu, Qingyang Li, Hui Xu
arXiv:2609.26443v1 Announce Type: new
Abstract: Recent advances in Artificial Intelligence (AI)-powered Computer-Aided Diagnosis (CAD) systems have substantially improved breast cancer screening, dia...
By Farnoush Bayatmakou, Maryam Hosseini, Reza Taleei, Arash Mohammadi
arXiv:2609.26733v1 Announce Type: new
Abstract: The rapidly advancing capabilities of vision-language models (VLMs) present a systemic challenge to academic integrity. VLMs now allow students to bypa...
By Christopher Burger, Christina Trotter, Joseph Carlisle, Charles Walter
arXiv:2609.25831v1 Announce Type: cross
Abstract: Recent VLM-based autonomous driving planners adopt GRPO-style reinforcement learning to optimize driving performance. However, existing GRPO recipes...
By Yuqi Ye, Shangkun Sun, Junhong Lin, Jiayi Zhao, Changhao Peng, Wei Zheng, Guoqing Liu, Tiesong Zhao, Wei Gao
arXiv:2609.25627v1 Announce Type: cross
Abstract: General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate prec...
By Haoran Wen, Wenfu Wang, Kunsong Shi, Jingke Wang, Wancheng Feng, Yiren Zhang, Yueran Zhao, Xuancheng Zhang, Nanfei Ye, Xingru Chen, Zhaohong Sun, Chengmin Yang, Zikang Yu, Penghao Bi, Jia Shi, Yu Liu, Kun Zhan, Yan Xie
arXiv:2512.02697v4 Announce Type: replace
Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images matching a query image. However, the traditional satellite-...
By Zixuan Song, Jing Zhang, Di Wang, Zhiming Luo, Wenbin Liu, Haonan Guo, En Wang, Bo Du, Liangpei Zhang
arXiv:2609.24526v2 Announce Type: replace
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
By Foundation Model, Li Auto Inc
arXiv:2607.04745v2 Announce Type: replace-cross
Abstract: Strong mapped-region thermal visual place recognition (VPR) does not ensure safe rejection of unmapped queries. We identify and quantify this...
By Zhiyuan Lu, Kanji Tanaka
RiverVLN introduces the first benchmark for long‑horizon vision‑language navigation (VLN) of unmanned surface vehicles (USVs) in continuous riverine motion. The PGT‑NAV framework converts navigation instructions into an ordered sequence of visually verifiable semantic phases, maintaining an active phase online through grounded visual and motion evidence. This phase‑grounded approach reduces recursive position and heading drift, achieving a 0.79 success rate in Unity‑ROS closed‑loop tests and demonstrating transfer to real‑world USV deployment.
By Jieling Wu, Yuehao Huang, Jiajun Lv, Tao Huang, Yong Liu, Weiwei Liu
The paper presents a practical approach to semi‑supervised federated learning for automatic speech recognition (ASR). It demonstrates that using a per‑client online teacher combined with a stabilizing server‑side anchor—where the server continues training on labeled data between rounds—significantly reduces divergence caused by pseudo‑label errors. The authors provide design guidelines that improve in‑domain performance by an average of 20.8 % and cross‑domain performance by 10.0 % over the best prior method, narrowing the gap to fully‑supervised federated learning.
By Wonho Bae, Zakaria Aldeneh, Martin Pelikan, Jan "Honza" Silovsky, Tatiana Likhomanenko, Sheikh Shams Azam
The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.
By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
The paper introduces modality‑gated deep adapters, a parameter‑efficient method for adding new modalities to a frozen multimodal embedding language model without altering its existing outputs. These adapters are bottleneck modules attached to each decoder layer, grouped into modality‑specific packs that activate only during encoding of their own modality, ensuring exact preservation of the base model’s computation graph. Experiments on a 2B base model show significant gains in audio‑to‑text and thermal‑to‑text retrieval metrics, and the authors release the audio and thermal packs along with training and evaluation code.
By Abdul Basit Tonmoy, Kazi Fardinul Hoque, Md. Shahrier Islam Arham, Arman Luthra