The study applies causal tracing to large audio language models (LALMs) to uncover how they fuse acoustic and textual information. Layer‑wise analysis reveals distinct fusion strategies—progressive integration in DeSTA versus abrupt late‑stage fusion in Qwen—while token‑wise analysis identifies the final sequence token as an informational bottleneck that decisively retrieves audio content. Additionally, an attention‑like query mechanism at intermediate tokens is observed, prompting the model to pull task‑relevant audio context.
By Wei-Chih Chen, Chien-yu Huang, Hung-yi Lee
AgroBench is a reproducible benchmark that converts U.S. county-level crop yield statistics into weakly supervised pixel‑level crop time series. The data generation pipeline fuses USDA yield data with land cover masks, Sentinel‑2 and Sentinel‑1 imagery, climatic variables, and terrain information to produce multimodal sequences for individual crop pixels across the growing season. The benchmark includes over 13 million observations from 788,654 crop pixels, covering 5,107 county‑year combinations for five major U.S. crops from 2017 to 2024, and establishes a Leave‑One‑Year‑Out evaluation protocol with baseline machine learning results.
By Udaiveer Singh, Rajiv Ranjan, Shashank Tamaskar, Dharmendra Saraswat
The paper introduces Motion Vision CAPTCHA (MVCAP), a new CAPTCHA framework that relies on motion-defined foreground structures to create challenges that are only solvable through temporal analysis of a dynamic background. MVCAP is implemented in three progressive levels—coherent motion, structural motion, and biological motion—and evaluated using the MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances. Human participants achieve 99.6% accuracy, whereas the best GUI agent scores only 16.8%, highlighting a significant human–agent perception gap and demonstrating that dynamic background camouflage is the key difficulty.
By Zeyu Zhang, Dingyi Rong, Zijian Chen, Zicheng Zhang, Xiongkuo Min, Guangtao Zhai
CereVLA is a cerebellum-inspired framework that enhances vision‑language‑action (VLA) policies by adding lightweight residual refinement and predictive consequence evaluation to frozen VLA execution. It generates corrective actions via flow‑based residual refinement, then assesses their short‑ and interval‑horizon impacts using a recurrent state‑space model and a history‑aware classifier, suppressing unfavorable corrections with a lightweight governor. Experiments on LIBERO‑10, LIBERO‑GOAL, and SO‑101 show that CereVLA improves task success rates and reduces control steps compared to state‑of‑the‑art baselines.
By Shuai Zeng, Yuxuan Liang, Hangmiao Hu, Fobao Zhou, Zixiang Wang, Wenxi Hong, Hang Zhao
AWM‑VLA introduces a unified framework that embeds aligned world modeling directly into a diffusion‑transformer vision‑language‑action policy. By adding learnable future tokens aligned with vision‑language embeddings of future observations, the policy can anticipate long‑term consequences while generating actions. The method extends this with an object‑centric alignment objective and a principled weighting scheme, achieving up to 21% higher success rates on RoboCasa and humanoid tabletop benchmarks and producing object‑centric rationales preferred by human raters in 83% of cases.
By An Lanji, Dawei Liu, Jin Li, Haoran Xu, Mei Chen, Yu Tian
DMM-Align introduces a closed‑loop framework for 2D‑3D registration that jointly refines correspondences, estimates pose, and learns representations using a shared differentiable geometric state. The method employs two diffusion processes: a geometry‑aware diffusion that improves the soft matching matrix for robust correspondence estimation, and a geometry‑conditioned diffusion teacher that feeds pose‑induced supervision back into feature learning. Experiments on 7‑Scenes and RGB‑D Scenes V2 show that DMM‑Align outperforms strong baselines, particularly in low‑overlap and heavily occluded scenarios, demonstrating the value of closed‑loop geometric feedback.
By Chongjian Wang, Junjie Gao
Groundbench is a new benchmark that evaluates vision‑language models on multi‑resolution polygon grounding, using the same 1,500 image‑expression‑referent triples but targeting exact‑N polygons with five different vertex budgets. It audits both filled‑region intersection‑over‑union (IoU) and legal‑polygon completion, revealing that the best models achieve 88.2 box IoU and 97.1 accuracy at IoU ≥ 0.5, while direct polygon predictions lag at 57.7 and 69.2. The study shows performance is non‑monotonic across budgets, collapses at the densest budget due to legality failures, and highlights that false spatial cues hurt more than false colour cues, underscoring an operational geometry gap beyond latent boundary perception.
By Zhonghan Bian, Zhenran Wang, Jinsong Li, Zhangyang Qi
UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.
By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.
By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
AstraLOD3 is a zero‑shot multimodal agentic system that reconstructs LOD3 building models using multi‑view images, calibrated cameras, a filtered sparse SfM point cloud, and a natural‑language specification. The Astra agent selects and executes computational steps in Python and Blender, achieving a mean FRDS of 0.9647 across 35 runs, including 24 benchmark buildings, with geometric agreement comparable to purpose‑built methods. Ablation studies show the impact of reconstruction guidance, evidence modalities, model configuration, and run‑to‑run variability, demonstrating that LOD3 reconstruction can be framed as a constrained agentic process rather than a fixed pipeline.
By Bryan G. Pantoja-Rosero
The paper introduces O$^2$-VG, a unified framework for oriented object visual grounding in remote sensing images, comprising three complementary models: O$^2$-VG-Trans, a cross‑modality transformer; O$^2$-VG-Uni, which predicts universal oriented proposals; and O$^2$-VG-VLM, an autoregressive vision‑language model that generates oriented bounding boxes. It also presents DIOR‑R‑SVG, a new dataset containing image, expression, and oriented box triplets for training and evaluation. The framework demonstrates superior performance across multiple benchmarks and is supported by publicly available code.
By Zeyu Ding, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Hancheng Zhu, Rui Yao, Abdulmotaleb El Saddik
Diff‑RF is a diffusion‑based framework that jointly performs image registration and fusion while accounting for degradation in multi‑modal images. It first restores modality‑specific degradations within each image, then uses a cross‑modal diffusion module that couples registration and fusion, refining alignment and enhancing complementary information. Experiments on extended datasets show that this coupled approach yields higher registration accuracy and fusion quality under diverse degraded conditions.
By Xunpeng Yi, Zaixi Du, Qinglong Yan, Yibing Zhang, Han Xu, Jiayi Ma
MultiVENT‑Raw is a new multilingual benchmark comprising nearly 120,000 raw videos—continuous footage from cell phones, hand‑held cameras, or CCTV—totaling over 5,300 hours. The dataset includes 130 events and 222 event‑centric queries, along with human‑annotated relevance judgments and extracted key facts for relevant videos. It supports two tasks: retrieving videos relevant to a query event and generating a coherent report summarizing event‑related videos for a target user, with baseline models showing these tasks remain challenging.
By Reno Kriz, David Etter, Alexander Martin, Cameron Carpenter, Debashish Chakraborty, Hannah Recknor, Reihaneh Iranmanesh, Matthew Maciejewski, Kenton Murray, Eugene Yang, Benjamin Van Durme, Aaron Steven White, Andrew Yates, William Walden
The paper introduces SGFNet, a Semantic‑Guided Fusion Network for classifying multi‑source remote sensing images. It features a Semantic Mixing Convolution Block that generates semantic‑aware kernels based on contextual relationships, and a Frequency Modulated Fusion Block that fuses cross‑modal information in the frequency domain to mitigate spatial misalignment. Experiments on the Augsburg and Houston 2018 datasets show SGFNet consistently outperforms state‑of‑the‑art methods.
By Yuwei Zhao, Chuanzheng Gong, Baogui Huan, Feng Gao, Junyu Dong, Qian Du
The paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task, aimed at interpreting complex tampering scenarios in land resource monitoring and national defense. It presents five datasets—RS-CMQA, RS-CMQA-B, Real-RSCM, RS-TQA, and RS-TQA-B—spanning 29 regions in 14 countries, and a region-discrimination-guided multimodal copy-move forgery perception framework (CMFPF) that improves question answering accuracy on tampered images. Experiments show CMFPF outperforms general VQA and RSVQA models, establishing a new benchmark for RSCMQA.
By Ze Zhang, Enyuan Zhao, Di Niu, Jie Nie, Xinyue Liang, Lei Huang
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
EnSIMem is an entity‑structured long‑term memory architecture designed for agents that interact with users over extended periods. It organizes interactions into theme‑coherent episodes and creates dialogue‑grounded index entries of the form [entity][entity type][property:value], preserving source turns, temporal data, and multimodal fields. During online interaction, the agent decomposes requests into evidence requirements, performs entity‑property lookup, and retrieves the necessary evidence to generate responses directly from preserved source material rather than lossy summaries.
By Xuanyu Meng, Xing Fan, Xinyi Fan, Chenlei Guo, Yixuan Xie, Jiawei Han
The study introduces SynerT, a waveform-only hybrid temporal model that uses a causal dilated TCN and dilated recurrent layers to predict early intraoperative acute kidney injury (AKI). Two extensions, SynerT-MM and SynerT-Stack, incorporate hemodynamic summaries, preoperative covariates, and a leakage-safe stacked ensemble to improve discrimination and calibration. Evaluated on the VitalDB database with a strict 60‑minute prediction window, SynerT-Stack achieved the best performance across AUROC, AUPRC, and F1‑max, and demonstrated the greatest net clinical benefit after recalibration.
By Quang Minh Nguyen, Duc Minh Le, Ho Nhat Minh Nguyen, Thuy Quynh Nguyen, Trong Nghia Nguyen
NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification shared task series and the second focused on multidialectal Arabic speech processing. It includes five main tasks—Automatic Speech Recognition, Spoken Dialect Identification, Text-to-Speech, Spoken Language Translation, and Spoken Language Understanding—along with eight subtasks that test realistic scenarios such as low‑bandwidth, mixed dialects, code‑switching, out‑of‑domain, and zero‑shot settings. The event attracted 21 teams from at least 13 countries, with 48 test‑phase submissions and 14 system‑description papers, and the results highlight out‑of‑domain generalization as a major bottleneck while showcasing the strengths of Arabic‑specialized speech models, multimodal dialect identification, and ensemble methods.
By Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch, Chiyu Zhang, AbdelRahim Elmadany, Youssef Mohamed, Salima Mdhaffar, Yannick Est\`eve, Mohamed Elhoseiny, Hamzah Luqman, Nizar Habash, Muhammad Abdul-Mageed
SpeechLLMs used in professional settings often undergo domain customisation through prompts or fine‑tuning, which can inadvertently cause the model to transcribe phonetically similar words from its context or training data, leaking private information. The authors systematically investigate this overlooked privacy risk, creating benchmarks to measure leakage rates for both prompting and fine‑tuning, and find that both mechanisms cause measurable leakage that compounds when combined. They evaluate a prompt‑level mitigation strategy and analyse the accuracy‑leakage trade‑off, concluding that fine‑tuning without context prompts offers the best balance between performance and privacy.
By Maike Z\"ufle, Jan Niehues