VISTA is a method for active multimodal agents that internalizes collective visual experience via on‑policy distillation. It turns observations from multiple rollouts of the same input into shared supervision, using Collective Visual Experience Distillation (CVED) to organize observations with context and Heterogeneity‑Aware Policy Improvement (HAPI) to reinforce successful trajectories and guide learning from unsuccessful ones. The approach lets an experience‑conditioned teacher evaluate a student’s partial responses, enabling discoveries from one trajectory to inform others without altering the student’s original history, and achieves superior performance on fine‑grained perception and general reasoning tasks compared to comparable agents.
By Zheng Jiang, Houde Qian, Yiming Chen, Ling Li, Chaoyang Li, Yueqi Li, Yuxuan Liu, Lifeng Sun
The paper introduces a controlled benchmark for evaluating large language models (LLMs) on key‑value pair extraction from documents with varying levels of OCR noise. It tests 136 configurations across five instruction‑tuned open‑weight LLMs, three datasets, and four text‑quality conditions, using deterministic decoding to generate 17,688 document‑level inferences. The study finds that clean‑text performance does not reliably predict real‑world robustness, model rankings can reverse under noisy conditions, and few‑shot demonstrations do not always improve accuracy, highlighting reliability risks in OCR‑to‑LLM pipelines.
By Zahra Anvari
Spatial-OPSD is a label‑free self‑improvement framework for vision‑language models that leverages spatial priors such as depth, 3D relations, and camera geometry to provide dense token‑level supervision. During training, a privileged teacher uses these priors while the student learns from only the original visual‑language input, and a recursive round‑wise scheme allows repeated self‑improvement without moving the teacher. Across four VLM families, one round of Spatial‑OPSD improves the five‑benchmark average, and three rounds push a strong spatially specialized model to the open‑source frontier, achieving the highest average among open models and best results on three of five spatial reasoning benchmarks.
By Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, Ruqi Huang
arXiv:2609.37024v1 Announce Type: new
Abstract: Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides pa...
By Junbo Shen, Jinying Gao, Bo Lei
arXiv:2609.37200v1 Announce Type: new
Abstract: Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary...
By Songlin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang, Toyota Li, Eric Liu, Alan Zhao, Anyi Rao
arXiv:2609.37314v1 Announce Type: new
Abstract: Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extre...
By Abhay Kumar Das, Rajesh Gangireddy, Ashwin Vaidya, Samet Akcay
arXiv:2609.37888v1 Announce Type: cross
Abstract: Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of visi...
By Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou
AnyAct introduces a universal action layer that consolidates diverse tool capabilities into a self‑evolving action space for AI agents operating in open‑world environments. It tackles the scale dilemma, tool non‑stationarity, and heterogeneous feedback by using hierarchical progressive retrieval and test‑time reliability evolution, while a heterogeneous observation grounding module unifies multi‑modal feedback. Evaluations on LiveMCPBench and the newly created OSMCP benchmark show state‑of‑the‑art performance, with significant gains in task success rate and reduced execution steps, especially for models with limited native capabilities.
By Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang
arXiv:2609.36934v1 Announce Type: new
Abstract: Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signal...
By Pan Zhang, Siqi Lai, Kemu Dong, Hao Liu
The paper introduces a curriculum reinforcement learning approach to overcome the cold‑start problem in prompt‑injection red‑teaming of frontier large language models. By training an attacker LLM sequentially against increasingly robust target models and ensuring partial success at each stage, the method achieves high attack success rates (93.8% against GPT‑5.6‑Luna and 45.0% against GPT‑5.6‑Terra) where prior RL methods fail. The attacker LLM also transfers its effectiveness to other frontier models it was not explicitly trained on.
By Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia
THEIA is a multimodal dataset that pairs thousands of analog circuit layout images with question‑answer conversations, and it introduces a benchmark using a fine‑tuned vision‑language model to analyze GDSII files. The dataset and benchmark enable designers to interact with and query physical layouts as intuitive, meaningful entities. Experiments on five realistic tasks show the fine‑tuned model outperforms general‑purpose vision‑language models by up to 73%, revealing a significant gap between general multimodal reasoning and domain‑specific layout understanding.
By Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni
The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.
By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
NeuronEye is a plug‑in framework that builds a sparse, concept‑level neuron vocabulary from intermediate vision‑language model (VLM) representations and selectively activates query‑relevant visual concepts during inference. It decomposes vision‑token states into an overcomplete sparse basis organized by concept clusters, uses the language query to activate relevant clusters, localizes the corresponding image patches, and injects the focused evidence back into the vision tokens, while a suppression mechanism attenuates dominant perceptual directions. Experiments on Qwen2.5‑VL‑7B and LLaVA‑1.6‑7B show that NeuronEye improves CV‑Bench overall accuracy by +3.1, boosts Distance by +9.5, and raises BLINK Multi‑view by +8.3, indicating that sparse neuron vocabularies can act as active interfaces for concept‑level visual reasoning.
By Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao
The paper investigates whether multimodal large language models (MLLMs) can generate and detect realistic multimodal fake news on social media. Using a multi‑agent framework—comprising a story agent, an image agent, and a critic agent—the authors produced over 9,000 paired multimodal news posts across science, health, and entertainment domains. They benchmarked 16 open‑ and closed‑source MLLMs for automated detection and found that most models fall far short of human accuracy, especially in identifying image authenticity, highlighting the need for stronger defenses against social media fake news.
By Jiyao Yang, Yang Liu, Zhenyue Qin, Qingyu Chen, Xiuzhen Zhang
arXiv:2609.32352v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging r...
By Gujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan, Yanlin Qi, Gangyi Zhang, Yuxin Yang, Dawei Li, Haiming Tang
GA-EIRFS is a detector‑agnostic sampling strategy that augments frequency‑based repeat‑factor sampling with a fixed geometry score derived from point count, surface‑normal entropy, and surface coverage. It only alters frame‑sampling probabilities, leaving the underlying detector and inference pipeline unchanged. Experiments on nuScenes show consistent improvements in mean average precision and the nuScenes detection score across multiple seeds and backbones, with notable gains for rare classes such as bicycles.
By Taufiq Ahmed, Constantino \'Alvarez Casado, Daniel Herrera Castro, Sasan Sharifipour, Abhishek Kumar, Miguel Bordallo L\'opez
LongLive‑Plug is a once‑for‑all distillation framework that learns reusable LoRA adapters on a base video diffusion model, enabling training‑free, plug‑and‑play deployment to a wide range of downstream models. These adapters provide single‑pass classifier‑free guidance, few‑step sampling, and long‑context error correction for autoregressive generation, and remain effective even when downstream models add conditioning branches or expand output channels. The authors demonstrate that the approach works on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation.
By Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
arXiv:2609.36440v1 Announce Type: new
Abstract: Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent...
By Jae-Ho Lee, Jeong-Eun Lee, Gyeong-Moon Park
arXiv:2609.36064v1 Announce Type: cross
Abstract: Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly...
By Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi, Jeffrey Landes, Angela Ju, Muhammad Azam, Bryan Lim, Johan Luttun, Indrajeet Haldar, Jonathan Shaw, Beatriz Guerra, Ivan Sosnovik, James Stoddart, Robert Giaquinto, Adam Gaier
arXiv:2609.37043v1 Announce Type: new
Abstract: Sampling from unnormalized distributions over large discrete state spaces becomes difficult when a multimodal target is far from a tractable reference....
By Yuwen Qian, Yidong Ouyang, Zhengyan Wan, Hongyuan Zha