The paper presents a mechanistic defense for Vision‑Language‑Action (VLA) models against adversarial patches. By using a sparse autoencoder, the authors identify a feature whose activation correlates strongly with the presence of an adversarial patch and suppress this feature only when a linear probe detects an attack. This conditional intervention improves robustness on the LIBERO‑10 benchmark while avoiding the performance degradation that occurs with continuous suppression.
By Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi
The paper introduces AdsCVR, a benchmark for e‑commerce cross‑video reasoning with 2,483 videos and 6,110 QA pairs across six reasoning dimensions. It proposes AdSeek, an agentic framework that actively selects visual and audio tools during multi‑turn exploration and uses an offline trajectory rectification mechanism to improve reinforcement learning. AdSeek achieves 74.30% accuracy on AdsCVR, outperforming its backbone by 27.90 percentage points and generalizes to the open‑domain CrossVid benchmark.
By Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng
The paper introduces a unified exponential framework for generalising K‑fold cross‑validation (CV) accuracy into conservative risk bounds, called gamma‑CUBV. It models dependence between folds via a joint sub‑Gaussian proxy matrix, yielding an effective number of folds and showing that more folds do not always increase evidence when data are strongly correlated. The framework extends to posterior predictor distributions and weighted multi‑source fusion, and is empirically evaluated on linear classifiers trained on heterogeneous multimodal Gaussian mixtures, demonstrating that corrected resubstitution can be tighter than K‑fold CV in low‑sample, high‑heterogeneity settings.
By JM Gorriz
TasteBench is a multimodal benchmark designed to accelerate sustainable protein discovery by providing computational proxies for sensory prediction. It includes a food-level ranking task based on over 21,000 human evaluations of 215 plant-based foods across 24 categories, and a molecular-level taste classification task covering 15,000 flavor molecules. The benchmark offers baseline models, characterizes inter-rater agreement and reliability limits, and demonstrates that the best model achieves pairwise accuracy comparable to individual human panelists.
By Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling
HakushoBench is a Japanese chart and table visual question answering benchmark created from 33 governmental white papers, comprising 2,053 images across more than ten types. The dataset includes manually annotated and independently verified QA pairs that evaluate holistic understanding of charts and tables rather than just local visual cues. Experiments show that HakushoBench is significantly harder than existing Japanese benchmarks, with sub‑10B open‑weight models achieving at most 58.6% accuracy and even large models like Qwen3.5-397B-A17B lagging behind Gemini‑3‑Pro by 8.1 points.
By Issa Sugiura, Shuhei Kurita, Yusuke Oda, Naoaki Okazaki
The paper introduces Function‑Space Guided Multimodal Optimization (FGMO) to address modality imbalance in multimodal learning. FGMO uses Functional Progress Estimation (FPE) to measure each modality’s function‑space response and Functional Response Control (FRC) to redistribute learning‑rate budgets, aiming to balance optimization progress across modalities. The authors provide theoretical analysis of FRC’s target‑contraction property and demonstrate FGMO’s effectiveness on several multimodal benchmarks.
By Zhongjing Gu, Fengqiang Wan, Yiming Cui, Yufa Feng, Yang Yang
The paper presents a geometry-driven framework for beam management in high-frequency massive MIMO systems. It constructs an offline virtual base station database using 3D LiDAR point clouds and location data to model dominant reflection paths, enabling a coarse channel reconstruction. A VBS-assisted orthogonal-pilot scheme and a dual-agent dueling double deep Q-network are then employed to refine beam estimates and perform coordinated beam selection, yielding improved training efficiency and performance over existing baselines.
By Yijie Bian, Wei Guo, Jie Yang, Shenghui Song, Jun Zhang, Shi Jin, Khaled B. Letaief
The paper introduces Multimodal Hypergraph Flow Matching (MHG‑FM), a framework that jointly generates structural connectivity (SC) and functional connectivity (FC) by constructing modality‑specific hypergraphs and learning higher‑order representations with Hypergraph Neural Network encoders. It employs Dual Cross‑Attention for bidirectional cross‑modal fusion, a variational autoencoder to map fused representations into a latent space, and conditional flow matching to synthesize connectivity and perform multimodal translation. Experiments on the Human Connectome Project Young Adult dataset demonstrate that MHG‑FM outperforms state‑of‑the‑art baselines in reconstruction quality, topology preservation, distributional similarity, and SC‑FC coupling, while achieving roughly eight times faster sampling than a comparable diffusion backbone.
By Chyong Yi Poh, Hwa Hui Tew, Junn Yong Loo, Rapha\"{e}l C. -W. Phan, Fuad Noman, Pew-Thian Yap, Chee-Ming Ting
The paper introduces Variational Risk Minimization (VRM), a distillation framework that treats large vision‑language model (LVLM) report variants as Monte Carlo samples of latent clinical interpretations. VRM learns from a variationally marginalized teacher distribution, providing uncertainty‑aware supervision even when modalities are missing, and outperforms fine‑tuning baselines while improving calibration. In a compact edge‑student setup, a confidence‑gated cascade achieves an AUC of 0.941 at 103 ms latency with only 20.3 % cloud escalation, offering a clear reliability‑latency trade‑off for cloud‑edge medical imaging workflows.
By Xinye Yang, Zhusi Zhong, Scott Collins, Michael Bernstein, Grayson Baird, Terrence Healey, Michael Atalay, Mahesh Jayaraman, Xuyu Wang, Zhicheng Jiao
Zephon is a data loader designed for foundation model training that ensures deterministic ordering of training data batches even when GPU resources change, checkpoints are resumed, or different processing backends are used. It handles online, stateful pipelines—where tokenization, packing, and mixing of samples create complex n‑to‑m transformations—by partitioning the data stream into topology‑independent lanes, serializing ordering decisions, and checkpointing only bounded in‑flight state. Experiments on text and vision‑language tasks show Zephon delivers competitive throughput while offering guarantees that existing loaders lack for such pipelines.
By Maximilian B\"other, Josh Wills, Ties Robroek, Sonnet Xu, Paul Burstein, Daniel Zayas, Cody Blakeney, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Haakon Mongstad, Luke Merrick, Pratyush Maini, Ari Morcos, Matthew Leavitt, Ana Klimovic, Bogdan Gaza
The paper proposes a method for image-based traversability estimation that uses von Mises‑Fisher mixture prototypes in a frozen vision‑language feature space. By incorporating relative natural‑language rules as a commonsense prior and fine‑tuning with sparse image annotations, the approach achieves domain adaptation with fewer labels. Experiments on WayFAST show competitive accuracy to end‑to‑end models and improved dense predictions on semantic maps, while qualitative results illustrate zero‑shot language prior use and semantic interpretation of prototypes.
By Simon Schwaiger, David Seyser, Alessandro Scherl, Zlatan Ajanovi\'c, Wilfried W\"ober, Gerald Steinbauer-Wagner
FiberGeoText (FGT) is a vision‑language model that clusters short‑range superficial white matter streamlines from ultra‑high‑resolution diffusion MRI into population‑level groups. It jointly encodes each streamline’s 3‑D trajectory, cortical anatomical context (via text from multiple parcellation schemes), and shape, using a pretrained large language model to unify heterogeneous anatomical descriptions. Evaluations on 0.76 mm diffusion data show that FGT outperforms state‑of‑the‑art methods in cortical parcel coherence, shape consistency, cluster‑size consistency, and cross‑subject correspondence, and it generalizes well to unseen subjects, recovering 96.7 % of learned clusters.
By Yuqian Chen, R. Jarrett Rushmore, Guikun Chen, Fan Zhang, Edward Yeterian, Nikos Makris, Yogesh Rathi, Lauren J. O'Donnell
The paper introduces Causal-Invariant Masking (CIM) to better quantify epistemic uncertainty in Multimodal Large Language Models (MLLMs) by measuring semantic shift between original predictions and those conditioned on a causally-focused view. It proposes Semantic Divergence as a core metric that converges to the variance of the model’s sensitivity to non‑causal correlations, and introduces Expected Embedding Drift (EED) as a fast geometric proxy that estimates this shift directly in the embedding space. Experiments demonstrate state‑of‑the‑art uncertainty quantification performance and a nearly 50% speedup with EED.
By Haoyang Luo, Linwei Tao, Jie Gui, Xinghao Chen, Chang Xu, Jianyuan Guo, Minjing Dong
The paper introduces Spatial Memory Intelligence (SMI), a framework that enhances long‑video world models by systematically managing spatial memory using an understanding model. SMI employs four coordinated operations—spatial clustering, within‑cluster sparsification, action‑aware retrieval, and reliability‑aware filtering—to handle increasingly complex and lengthy memory sequences. Experiments across various baselines and benchmarks show that SMI improves memory sparsity, generation stability, and spatial consistency, demonstrating its effectiveness and generalizability.
By Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang
AMBER is a new online, budgeted multi‑view reranking framework for vision‑language models that dynamically allocates computation to maximize information gain. It treats fragmented listwise VLM outputs as local tournaments and uses continuous Elo updates to maintain a lightweight global ranking state. Experiments on CIRR, CIRCO, and PhotoBench show that AMBER outperforms other multi‑call VLM reranking methods under comparable budgets, and remains effective even with lower budgets.
By Wenteng Chen, Jiachen Zhu, Rong Shan, Tianyi Xu, Yuxiang Chen, Congmin Zheng, Teng Wang, Junjie Wu, Weiwen Liu, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
World Action Planner is a robot planning system that uses an action-conditioned world model to search for and compose executable action plans. The agent performs a coarse-to-fine search: first a global action optimization over imagined rollouts to spot potential failures, then a local action search comparing neighboring candidates to pick the best action. In compositional long-horizon tasks, novel object layouts, and real‑robot planning without expert demonstrations, it consistently outperforms state‑of‑the‑art end‑to‑end generalist policies and VLM planners.
By Xiangcheng Zhang, Runhan Huang, Yilun Du
The paper proposes a scalable traffic modeling approach that uses a single representative large language model (LLM) agent for each homogeneous traveler group, rather than one LLM per traveler. The representative agent maintains a mixed strategy over routes, updates it daily based on positive reinforcement signals, and uses a tunable step size to adjust its strategy. This design improves scalability, stabilizes learning, and produces interpretable dynamics that reproduce realistic behavioral patterns such as the decoy effect and income‑based willingness‑to‑pay differences.
By Hanlin Sun, Jiayang Li
OpenBox is a two‑stage automatic annotation pipeline that uses a 2D vision foundation model to associate 2D image cues with 3D point clouds. It then classifies instances by rigidity and motion state to generate adaptive bounding boxes using class‑specific size statistics, eliminating the need for self‑training. Experiments on Waymo Open Dataset, Lyft Level 5 Perception, and nuScenes show improved accuracy and efficiency over existing baselines.
By In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
The paper introduces SafeCut, a method for Source-Free Domain Adaptation that uses Vision‑Language models as external knowledge. SafeCut employs a cut statistic to gauge prediction reliability, enabling a dynamic, reliability‑gated supervision between the source‑pretrained model and the ViL model. This approach selectively amplifies correct mutual corrections while suppressing error propagation, achieving state‑of‑the‑art performance on multiple SFDA benchmarks.
By Seongjun Lee, Changhee Lee
VisAudit is a new benchmark that tests multimodal agents on visual diagnosis, repair, and verification tasks. It presents agents with rendered charts and auxiliary evidence—such as source data, intended summaries, and code—to iteratively detect defects, modify the visualization, and confirm successful repairs. The benchmark includes 1,900 flawed charts across 21 types and 10 flaw categories, plus 300 correct charts, and shows that current models recover only about 47.4% of flawed charts autonomously.
By Shicheng Liu, Adam Kahirov, Qi Zhang, Zhimin Hu, Song Wang, Junhong Lin, Julian Shun, Yada Zhu