The paper introduces Spatial Memory Intelligence (SMI), a framework that enhances long‑video world models by systematically managing spatial memory using an understanding model. SMI employs four coordinated operations—spatial clustering, within‑cluster sparsification, action‑aware retrieval, and reliability‑aware filtering—to handle increasingly complex and lengthy memory sequences. Experiments across various baselines and benchmarks show that SMI improves memory sparsity, generation stability, and spatial consistency, demonstrating its effectiveness and generalizability.
By Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang
AMBER is a new online, budgeted multi‑view reranking framework for vision‑language models that dynamically allocates computation to maximize information gain. It treats fragmented listwise VLM outputs as local tournaments and uses continuous Elo updates to maintain a lightweight global ranking state. Experiments on CIRR, CIRCO, and PhotoBench show that AMBER outperforms other multi‑call VLM reranking methods under comparable budgets, and remains effective even with lower budgets.
By Wenteng Chen, Jiachen Zhu, Rong Shan, Tianyi Xu, Yuxiang Chen, Congmin Zheng, Teng Wang, Junjie Wu, Weiwen Liu, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
World Action Planner is a robot planning system that uses an action-conditioned world model to search for and compose executable action plans. The agent performs a coarse-to-fine search: first a global action optimization over imagined rollouts to spot potential failures, then a local action search comparing neighboring candidates to pick the best action. In compositional long-horizon tasks, novel object layouts, and real‑robot planning without expert demonstrations, it consistently outperforms state‑of‑the‑art end‑to‑end generalist policies and VLM planners.
By Xiangcheng Zhang, Runhan Huang, Yilun Du
The paper proposes a scalable traffic modeling approach that uses a single representative large language model (LLM) agent for each homogeneous traveler group, rather than one LLM per traveler. The representative agent maintains a mixed strategy over routes, updates it daily based on positive reinforcement signals, and uses a tunable step size to adjust its strategy. This design improves scalability, stabilizes learning, and produces interpretable dynamics that reproduce realistic behavioral patterns such as the decoy effect and income‑based willingness‑to‑pay differences.
By Hanlin Sun, Jiayang Li
OpenBox is a two‑stage automatic annotation pipeline that uses a 2D vision foundation model to associate 2D image cues with 3D point clouds. It then classifies instances by rigidity and motion state to generate adaptive bounding boxes using class‑specific size statistics, eliminating the need for self‑training. Experiments on Waymo Open Dataset, Lyft Level 5 Perception, and nuScenes show improved accuracy and efficiency over existing baselines.
By In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
The paper introduces SafeCut, a method for Source-Free Domain Adaptation that uses Vision‑Language models as external knowledge. SafeCut employs a cut statistic to gauge prediction reliability, enabling a dynamic, reliability‑gated supervision between the source‑pretrained model and the ViL model. This approach selectively amplifies correct mutual corrections while suppressing error propagation, achieving state‑of‑the‑art performance on multiple SFDA benchmarks.
By Seongjun Lee, Changhee Lee
VisAudit is a new benchmark that tests multimodal agents on visual diagnosis, repair, and verification tasks. It presents agents with rendered charts and auxiliary evidence—such as source data, intended summaries, and code—to iteratively detect defects, modify the visualization, and confirm successful repairs. The benchmark includes 1,900 flawed charts across 21 types and 10 flaw categories, plus 300 correct charts, and shows that current models recover only about 47.4% of flawed charts autonomously.
By Shicheng Liu, Adam Kahirov, Qi Zhang, Zhimin Hu, Song Wang, Junhong Lin, Julian Shun, Yada Zhu
The paper investigates the gap between retrieval and reading in document vision‑language models, showing that even when the correct page is retrieved, the model often fails to use the text on that page. By comparing answers derived from page images alone versus images plus extracted OCR text, the authors demonstrate that adding OCR text can improve strict accuracy by 13 to 16 points on their FoveDoc‑Bench benchmark. The study also reveals that OCR benefits textual evidence but not charts or figures, and that its advantage diminishes as retrieval quality worsens.
By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu
The paper evaluates how well generalist and dermatology-specific machine learning models perform on diverse skin lesion datasets, including dermoscopic images and smartphone photographs. It benchmarks a range of architectures—general-purpose vision-language models, foundation models, and task-specific dermatology classifiers—under conditions of distribution shift, modality change, and demographic variability. The study quantifies the performance gap between current state‑of‑the‑art models and the robustness needed for safe, equitable clinical deployment.
By Emanoel dos Santos, Kelvin Cunha, Rodrigo Mota, Fabio Papais, Thales Bezerra, Natalia Lopes, Erico Medeiros, Shirley Cruz, Jessica Araujo, Paulo Borba, Tsang Ing Ren
HyperBrowseComp is a multilingual and multimodal web‑browsing benchmark featuring 423 manually authored, human‑validated questions in 13 languages. The questions are intentionally difficult, requiring users to locate obscure evidence, follow multi‑step clue chains, or inspect heterogeneous sources such as videos, scanned documents, images, or maps. The benchmark filters out easier questions by testing models without internet access and evaluates performance using provider‑native search and a shared external retrieval harness, with a human evaluation on a sample to contextualize model effort.
By Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina
The paper introduces SCOUT, an offline multi-agent reinforcement learning framework that combines a generative behavioral prior with a decomposed value function for test-time action refinement. SCOUT uses optimal unified transport and Stein variational gradient descent to steer behavioral samples toward high-value regions, with the number of transport steps providing adaptive scaling. The authors prove a single-term KL bound on the joint soft-value gap under the individual-global-max principle and demonstrate that SCOUT outperforms existing methods on both discrete and continuous offline MARL benchmarks, including offline-to-online settings.
By Dongsu Lee, Haoran Xu, Amy Zhang
The study investigates why multimodal large language models (VLMs) perform better at verifying scientific claims when evidence is presented as a table rather than a chart, despite both formats containing the same data. Using layer‑wise linear probing and attention analysis on three open‑weight VLMs, the authors find that chart information is indeed encoded in intermediate representations but never reaches the prediction layer, a gap absent for tables. Attention patterns reveal that this disconnect manifests differently across model families, suggesting the issue lies in how encoded visual data is utilized at prediction time rather than in the encoding process itself.
By Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin, Akiko Aizawa
The paper demonstrates that a targeted adversarial perturbation can reduce a vision‑language model’s training loss to near zero for a fixed target caption, yet the same model, when generating freely, still produces the correct description. This phenomenon, termed the train/inference gap, is traced to a single autoregressive step where the target token’s rank is fixed across all images, and further analysis shows that the language decoder, rather than the visual encoder, determines whether the corrupted signal is amplified or suppressed. The study uses a controlled two‑stage PGD attack on Qwen2.5‑VL‑7B‑Instruct and evaluates the effect on 200 held‑out COCO images, revealing that adversarial robustness in autoregressive VLMs largely depends on the language decoder’s prior.
whyItMatters":"The findings suggest that defenses and faithfulness evaluations for deployed vision‑language models should focus on the language decoder rather than the visual encoder, as the former is the key determinant of robustness to adversarial perturbations."
By Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama
The paper introduces a unified framework for multimodal 3D human pose estimation that fuses RGB, LiDAR, and mmWave radar data while incorporating kinematics-based sensor fusion. It presents a black-box subject membership inference attack and a pointwise maximal leakage analysis to assess privacy risks, and proposes a user-level differential privacy method called Action Temporal Stratification to mitigate these risks. The framework is evaluated on the MM-Fi dataset under three experimental protocols, with source code to be released upon acceptance.
By Kaushik Bhargav Sivangi, Fani Deligianni
Multimodal representation learning has largely relied on contrastive models like CLIP that produce a single embedding per sample, which limits their ability to capture relation-dependent relevance. The proposed Relation-Conditioned Multimodal Learning (RCML) framework explicitly conditions embeddings on natural‑language relation descriptions, enabling the same sample to be represented differently under various relational contexts. RCML builds relation‑aware training pairs, incorporates a relation‑conditioned module, and uses a unified contrastive objective to jointly model cross‑modal alignment and relation‑induced structure, achieving superior performance on retrieval and classification tasks across zero‑shot, fine‑tuned, and out‑of‑domain settings.
By Yang Qiao, Yuntong Hu, Bowen Zhu, Hasibul Haque, Liang Zhao
The paper investigates how much information can be extracted from different internal representations of vision‑language models, focusing on two bottlenecks: low‑dimensional projections of the residual stream (via tuned lenses) and the final top‑k logits. It systematically compares the amount of retained information at these levels and finds that even the easily accessible top‑logit bottleneck can leak task‑irrelevant details from image queries, sometimes matching the leakage seen from full residual projections.
By Masha Fedzechkina, Eleonora Gualdoni, Rita Ramos, Sinead Williamson
The paper introduces an automated pipeline that reconstructs editable 3D procedural models of field‑grown maize directly from raw 3D point clouds, eliminating the need for manual tuning or species‑specific training data. It uses a vision‑language model to annotate leaf midlines in rendered views, then applies deterministic geometric algorithms and differentiable NURBS fitting to generate accurate plant descriptors and refine leaf surfaces. The method achieves a median Chamfer distance of 5.4 mm on 100 diverse maize plants and recovers 99.4% of reference leaves with high overlap, outperforming previous semi‑automated approaches.
By Mozhgan Hadadi, Talukder Z. Jubery, Adarsh Krishnamurthy, Baskar Ganapathysubramanian
DEPICT is a new training‑free metric for evaluating text‑to‑image alignment. It replaces fixed reference answers with an agreement rule that compares image‑based and caption‑only responses, weighting questions by how decisively the caption determines them. By merging this agreement score with a holistic score, DEPICT improves negation accuracy dramatically and outperforms existing training‑free metrics while matching or exceeding fine‑tuned evaluators on several benchmarks.
By Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
The paper investigates how stylistic changes introduced by large language models (LLMs) affect multimodal claim verification, a task that determines whether a textual claim is supported by given evidence. Two rewriting strategies are used: natural rewriting, mimicking typical academic polishing, and controlled injection, adding a single LLM-associated word. Across 11 open‑weight models (2B–38B parameters) from five VLM families, the study finds that most models remain robust to these modifications, showing no significant accuracy drop, though consistent probability shifts—especially under hedging conditions—are observed.
By Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa
The paper evaluates Vision Language Models (VLMs) on Visual Question Answering tasks where questions violate Grice's maxims. By generating question modifiers that add non-essential, ambiguous, or false information, the authors show that VLMs such as ChatGPT, Claude, Gemini, and Llava exhibit reduced performance. They also compare human pragmatic reasoning to VLM reasoning, noting differences in how each handles human‑induced versus AI‑generated violations, and find that humans spend less time resolving VLM‑induced violations while VLMs are less accurate in those cases.
By Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal