SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.
By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
ChartRevise is a new dataset and evaluation protocol designed for exact chart editing via code. It contains 92,438 records covering 344 edit types across 20 chart types and three plotting libraries, built using the grammar of graphics and source‑program checks to ensure applicability. The protocol measures atomic requirement completion, detects gratuitous changes and missed coupled updates, and combines these with execution and rendering success to determine exact‑edit success.
By Jiaxiang Tang, Yi Zhou, Chad DeLuca, Rogerio Feris, Ahmed Khalil Omran, Zhi-Li Zhang, Pengyuan Li, Ali Anwar
UniEvo‑VL is a self‑evolving framework that lets multimodal models improve themselves by using their own critiques as privileged information. The method trains a single model to act as both teacher and student, minimizing divergence between their diffusion distributions over sampling trajectories. Experiments on Qwen‑image‑2512 show significant gains in image generation metrics, and stronger external critics further raise the improvement ceiling, though results vary across tasks.
By Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi
The paper introduces Retrospective World Modeling, a new paradigm for vision‑language‑model (VLM) agents that allows them to reason backward by estimating which action most likely caused a state transition. It proposes the Self‑Consistency Reward (SCR), an intrinsic signal that measures how well a policy action aligns with this retrospective explanation, providing dense transition‑level feedback. Experiments demonstrate that incorporating SCR improves policy robustness and generalization compared to purely prospective world‑modeling approaches.
By Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo
DrivingBench is the first benchmark that tests general‑purpose vision‑language models on the task of driving a real Toyota Corolla around a parking‑lot cone course. The models receive live camera frames and issue steering and velocity commands, with inference latency counted as part of the challenge. In tests, only GPT‑6 Astra completed the course, while other models showed limited progress or failed to pass half the course.
By Aditya Ramabadran, Simon Mahns, Tobias Gessler
The paper introduces TED (Text-Axis Evidence Decomposition), a post‑hoc scoring method that improves anomaly localization in CLIP‑based detectors without altering the backbone or prompts. TED evaluates whether ambiguous responses are better supported by defect patches or normal patches, thereby distinguishing true defects from visually complex normal regions. Experiments show that TED significantly enhances pixel‑level localization across frozen VLM backbones and adapted hosts, especially under hard‑false‑positive competition.
By JinYoung Kim, Geonho Kim, GiJeong Park, Geonu Lee, YoungJoon Yoo
The paper introduces an LLM-based framework for continuous dimensional emotion evaluation in multimodal dialogue, combining discrete emotion recognition with Valence-Arousal-Dominance (VAD) assessment on the IEMOCAP dataset. It incorporates acoustic cues as natural language descriptions via the SpeechCueLLM approach and evaluates six models from the LLaMA, GPT, and Qwen families using zero-shot, few-shot, and LoRA fine-tuning. LoRA-fine-tuned LLaMA models outperform prompt-engineered GPT models, achieving a new state-of-the-art Valence CCC of 0.7822, and ablation studies show that textual audio descriptions significantly benefit smaller models.
"whyItMatters":"The study demonstrates that domain adaptation through fine-tuning can surpass larger GPT models in multimodal emotion evaluation, highlighting the importance of tailored training for emotion recognition tasks."
By Yutong Hu, Jinho Choi
CAMOS is a coupled oscillatory state‑space model designed for multimodal clinical time‑series that are irregularly sampled and often incomplete. It addresses a representational limitation of linear state‑space models by allowing the transition operator to depend on which modalities are present, using a bank of second‑order oscillators coupled through a gated matrix. On the ADNI dataset, CAMOS outperforms both uncoupled oscillatory models and clinical fusion models in same‑visit staging, landmark prediction, and longitudinal forecasting, and uniquely avoids collapsing to the majority class when transferred zero‑shot to OASIS‑3.
By Maxx Richard Rahman, Mostafa Hammouda, Wolfgang Maass
OmniReasoning introduces a new benchmark, OmniReasoningBench, that requires both audio and visual evidence for answering 1,150 multiple-choice and open-ended questions across two tasks. The authors also develop OmniQA, a data engine that automatically generates evidence‑grounded QA pairs with time‑stamped clue chains, producing training datasets OmniReasoning‑SFT‑112K and OmniReasoning‑RL‑19K. Finally, they propose Modality‑Factored Self‑Distillation (MFSD), an on‑policy self‑distillation method that assigns token‑level credit by evaluating responses under modality‑specific clue contexts, enabling the OmniReasoning‑30B‑A3B model to achieve significant performance gains on both the new benchmark and existing video benchmarks.
By Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu, Yiwu Zhong
GroundingPI is a 4‑billion‑parameter grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. It is trained with multimodal and spatial pretraining, supervised fine‑tuning, and reinforcement learning, achieving a new state‑of‑the‑art average of 73.68% across 34 grounding benchmarks. As a visual backbone, GroundingPI improves performance in robotic manipulation and autonomous driving, outperforming larger models and mainstream backbones in several out‑of‑distribution settings.
By Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
D‑Scope is a framework that links the interpretation of sparse autoencoder (SAE) features in diffusion transformers (DiTs) to controllable image generation. It aggregates SigLIP‑2 embeddings of highly activating image patches into visual centroids, matches target text descriptions against these centroids, and retrieves individual features without per‑feature text annotations. The method provides visual evidence for each selection and uses spatially masked interventions to test decoder directions under fixed generation conditions, evaluated across 150 SAEs and a benchmark of 100 target concepts.
By Xinyue Xu, Jiahao Zhang, Lijie Hu, Peter Hase, Hao Wang
The paper introduces FailBank, a four‑stage self‑evolving framework that transforms runtime feedback from safety shields into lasting policy improvements for vision‑language‑action (VLA) models. By using a counterfactual correction teacher, outcome‑aware admission, and guarded LoRA updates, FailBank converts useful shield proposals into corrective targets while preserving successful actions as anchors. Experiments on the VLA‑Arena benchmark show that FailBank boosts task success rates by up to 8.5 percentage points and reduces cumulative policy cost by up to 35.6%, outperforming both base policies and traditional runtime shielding.
By Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang
LEAP is a framework for long audio‑video question answering that avoids encoding entire recordings by dividing them into fixed‑duration blocks. It performs a lightweight localization pass on each block to score short candidate windows, then pools the highest‑ranked windows for a single bounded answer pass, keeping the answer input and peak context independent of recording length. The method trains both a localization LoRA and an answer LoRA, supports causal streaming queries, and achieves significant performance gains over baseline models on multiple AVQA benchmarks.
By Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng
GateSPINE is a vision‑language framework designed for automated lumbar spine MRI report generation. It fuses sagittal T1 and T2 volumes using a training‑free gated cross‑view fusion module, then encodes the fused sagittal and axial volumes with parallel 3D encoders before decoding the combined representation into a report. Evaluated on three datasets, GateSPINE achieves the highest clinical efficacy F1 scores, improving recall across all datasets while remaining competitive on standard natural language generation metrics.
By Hoang Nguyen Van, Cuong Vuong Tuan, Trang Mai Xuan, Bien Tran Van, Nam Tran Van, Thien Van Luong
arXiv:2605.09860v5 Announce Type: replace
Abstract: Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning. This number, the execu...
By Chen Li, Zhantao Yang, Fangyi Chen, Han Zhang, Anudeepsekhar Bolimera, Marios Savvides
arXiv:2609.25187v2 Announce Type: replace
Abstract: Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems...
By Howard Lu, Shalfun Li, Porter Pan, Cris, Lumen, Cyril, Eric Hu, Lily Li, Maeve Zhang, Rain Sun, Robert Wang, KZ Zheng, Viggo Chen, Tim Ding, Regsis Cheng, YJ Xiao, Kian, Hai Lin, Alan Song, Elise Ma, Gody Li, Victor Yao, Yohann Tang, Ingrid Yu, Jason He, James Wang, Ryan Yu, Ping Yang, Chris Pan, Vincent Chen, Roy Gan, Hao Wang, Qian Wang
arXiv:2609.34211v2 Announce Type: replace
Abstract: In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However,...
By Linbo Shao, Huilin He, Yating Lou, Dawei Cheng
arXiv:2510.17932v5 Announce Type: replace-cross
Abstract: We introduce Chart2Code, a new benchmark for evaluating the chart understanding and code generation capabilities of large multimodal models (...
By Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang
arXiv:2602.08005v2 Announce Type: replace-cross
Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computati...
By Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu
arXiv:2603.10210v2 Announce Type: replace-cross
Abstract: While Diffusion Models excel in text-to-image synthesis, they frequently suffer from catastrophic concept omission when generating complex mu...
By Zitong Wang, Zijun Shen, Haohao Xu, Zhengjie Luo, Weibin Wu