Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context...
Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typica...
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tr...
Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relyi...
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities inde...
This work proposes a mutual feedback architecture, MEQ, that refines the two inputs, of possibly different modalities, into a pair of coupled embeddings such that each embedding reflects the informati...
Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collaps...
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and pr...
Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information the question needs, or because the solver fails to...
CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase...
Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector and LoRA adapters of a speech LLM at five audio com...
CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.
By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.
By Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
PARSEE-VAD is a training‑free online video anomaly detection framework that separates semantic evidence acquisition from score‑state evolution. It uses Proposition‑Aware Reasoning to extract structured propositional evidence from the current causal window and selectively activates more specific queries, while Streaming Evidence Escalation maps this evidence into a compact score‑domain event state and propagates only the bounded state to maintain temporal continuity. Experiments on four benchmarks show strong performance with reduced specialist computation and sparse score‑state propagation, supporting a current‑first principle for streaming multimodal inference.
By Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao
RawVLA introduces a streaming neural image signal processor that adaptively renders RAW observations for vision‑language‑action (VLA) policies, focusing on imaging factors that influence embodied behavior. The authors systematically analyze how five ISP dimensions—gain, sensor noise, chromatic response, tonal response, and bit depth—affect action prediction and manipulation success, showing that RAW‑to‑RGB processing significantly shapes outcomes. They also present RawVLA‑Bench, a RAW‑domain manipulation benchmark that evaluates image processing as an explicit variable across clean and adverse conditions, demonstrating that RawVLA maintains performance under standard settings while markedly improving robustness under degraded imaging.
By Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui
The study investigates whether a prosody‑trained representation can improve automatic speech recognition beyond the effect of a trainable fusion mechanism. Using a frozen HuBERT backbone and a 64‑dimensional prosodic representation, the authors compare a baseline recognizer, a fusion model with no auxiliary input, and a fusion model with the learned representation across Buckeye, Switchboard, and AMI IHM datasets. While the fusion model without auxiliary input reduces WER relative to the baseline, adding the prosodic representation yields no significant WER improvement, though the model still depends on the representation for optimal performance.
By Ki Woong Moon, Daniel Brenner
ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.
By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang
InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.
By Hongpei Zheng, Hujun Yin
HyperSAM is a promptable foundation model for hyperspectral remote sensing that integrates a data‑centric synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). The model generates full‑spectrum hyperspectral cubes from high‑resolution multispectral imagery using a physics‑informed abundance‑transfer generator, and employs SAM3‑derived pseudo‑masks for object‑centric supervision. With a frozen SAM3 RGB branch, a trainable hyperspectral encoder, ControlNet‑style feature injection, and a mixture‑of‑experts mask refiner, HyperSAM demonstrates strong generalization across diverse hyperspectral tasks such as classification, anomaly detection, change detection, target detection, and airborne oil‑spill mapping.
By Li Pang, Xinqiao Wu, Jing Yao, Pedram Ghamisi, Jun Zhou, Zhengchao Chen, Deyu Meng, Xiangyong Cao
OmniVCBench is a figure‑centric, source‑traceable benchmark designed to evaluate the interpretation component of Artificial Intelligence Virtual Cells (AIVCs). It comprises 6,077 curated question–answer pairs drawn from scientific figures and experimental contexts, organized into three scientific reasoning tasks that mirror the AIVC Predict–Explain–Discover agenda. The benchmark also introduces AIVC‑Judge, a task‑conditioned MLLM‑as‑a‑judge framework with reference‑aware rubrics, and a Model‑Derived Hard‑Negative Mining strategy to generate multiple‑choice distractors for efficient evaluation.
By Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang