The paper introduces Rita, a reinforcement learning framework that addresses thinking drift in vision‑language models by enforcing consistency between reasoning and answers. Rita employs two reasoning‑label‑free rewards—thinking and consistency rewards—derived from the conditional probability of reference answers, and uses a difficulty‑aware data filtering strategy to select informative samples for training. Experiments on EgoIntention and RefEgo‑Int benchmarks demonstrate that Rita outperforms both supervised fine‑tuning and vanilla RL‑fine‑tuned approaches.
By Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao
The paper introduces TSDA-Track, a Template-Search Domain Adaptation framework designed to reduce modality gaps in cross‑modal visual object tracking. Two variants are explored: Pre‑AFA TSDA‑Track uses adversarial alignment before transformer interaction, while Enc‑CFA TSDA‑Track applies contrastive alignment after interaction to strengthen cross‑modal correspondence. Experiments on datasets such as LasHeR, RGBT234, GTOT, and Anti‑UAV‑024 show that both variants outperform state‑of‑the‑art trackers, with Pre‑AFA achieving an SR/PR of 43.2/56.0 on RGBT234 under the modality‑switch protocol.
By Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran
Contrastive On-Policy Distillation (COPD) is a framework that improves on-policy distillation by using a frozen teacher to evaluate student states under two contrasting prompts—one encouraging low reasoning effort and one encouraging high effort. The difference in log‑probabilities between these prompts provides a token‑level advantage signal that guides the student toward more concise and efficient reasoning strategies. Experiments on nine multimodal benchmarks show that COPD reduces reasoning length while maintaining task performance, and the contrastive approach can also be applied to on‑policy self‑distillation, allowing a model to compress its own reasoning without an external teacher.
By Jiacheng Ruan, Jun Tang, Wenzhen Yuan, Ting Liu, Shuai Bai, Dayiheng Liu, Zhibo Yang, Yuzhuo Fu
FoCLIP is a framework that creates adversarial examples to manipulate CLIP-based image quality metrics by reducing the alignment between image and text features. It uses stochastic gradient descent to combine feature alignment, score distribution balancing, and pixel‑guard regularization, enabling high CLIPscore predictions while maintaining visual fidelity. Experiments on artistic prompts and ImageNet show significant CLIPscore gains, and the authors also propose a color‑channel sensitivity detection method that achieves 91% accuracy.
By Yulin Chen, Zeyuan Wang, Tianyuan Yu, Yingmei Wei, Liang Bai
Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving proposes EMPlan, a hybrid trajectory planning method that combines sparse anchors with an offset refinement module for low-latency, high-accuracy predictions. The approach uses a two-stage training paradigm—pretraining followed by reward-guided fine-tuning—to improve safety without extra inference cost, leveraging rule-based reward signals and unpaired preference supervision. EMPlan is evaluated on the non-reactive NAVSIM benchmark, achieving a favorable balance between planning accuracy and efficiency under real-time constraints.
By Chenglin Chen, Lujia Wang, Xinhu Zheng, Jun Ma, Haoang Li
EviRover is a perception agent that goes beyond a single glance by actively gathering information to resolve perceptual queries. The authors created two data generation pipelines, producing EviRover-SFT-5K and EviRover-RL-12K, and a human‑verified benchmark called EviLens with 688 instances across five perception categories. Trained with supervised fine‑tuning and agentic reinforcement learning, the 4B EviRover outperforms its backbone by an average of 30 points on EviLens and shows strong transfer to other benchmarks such as WebEyes and BrowseComp‑VL.
By Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
The article discusses a notable imbalance in AI security research, where studies on attacking AI systems outnumber those on defending them. It highlights that this skew is evident across various subfields such as federated learning, speech recognition, membership inference, and large language models. The authors argue that attack papers often benefit from favorable evaluation conditions, whereas defense papers face stricter standards, resulting in a literature rich in vulnerabilities but lacking robust, deployable protections.
By Youqian Zhang
GATE-ST is a gene-aware text-image encoder that enhances spatial transcriptomics predictions by integrating gene descriptions into image-based models. The method encodes gene summaries with a text encoder and fuses these embeddings with image features via cross‑attention, aligning them with morphological cues. Benchmarks show GATE‑ST outperforms random gene embeddings and other image‑text fusion architectures, indicating its potential to improve accuracy while reducing time and cost in spatial gene expression analysis.
By Lucas Ni, Jian Luo, Wentao Huang, Chao Chen
The study investigates how humans and AI collaborate on a puzzle task, focusing on referential uncertainty—when a description could refer to multiple objects. It finds that eliciting a belief distribution over candidate pieces yields better calibration and discrimination than raw action probabilities, and that precise descriptions or well‑targeted hedges significantly reduce the acceptance of wrong placements. However, the AI rarely externalizes uncertainty, and poorly targeted hedges can be counterproductive.
By Christian Poelitz, Finale Doshi-Velez, Si\^an Lindley
arXiv:2609.38362v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining dis...
By Hung-Jen Chen, Yu-Heng Ho, Ting-Yao Huang, Po-Hsiang Hsu, Li-Yu Chen, Chun-Yi Lee, Min Sun
arXiv:2609.40014v1 Announce Type: new
Abstract: Detecting violence after it begins is important from recognizing behavioral cues that appear immediately beforehand. This work studies short-horizon pr...
By Sindhuja Penchala, Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi
arXiv:2508.16644v5 Announce Type: replace
Abstract: Diffusion models excel at photorealistic synthesis but struggle with object count fidelity, especially in high-density settings. We introduce COUNT...
By Anindya Mondal, Sauradip Nag, Ayan Banerjee, Josep Llados, Xiatian Zhu, Anjan Dutta
arXiv:2609.38897v1 Announce Type: cross
Abstract: Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emp...
By Shivam Saini, Eric Bezzam, Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen
arXiv:2609.39969v1 Announce Type: cross
Abstract: Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding tr...
By Yiming Gao, Shaocheng Luo
arXiv:2609.40325v1 Announce Type: new
Abstract: As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anoma...
By Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang
arXiv:2609.38397v1 Announce Type: new
Abstract: Virtual clients offer a cost-effective approach to support applications such as A/B testing, recommender system development, and interface evaluation....
By Yunan Lu, Shuang Xie, Meghna Allamudi, Mingyu Zhao, Han Li, Lingyun Wang, Zhou Yu
arXiv:2609.38755v1 Announce Type: new
Abstract: A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and rece...
By Zhining Gu, Shangjie Du, Weimin Qiu, Carl Olsson, Ping Liu, Meng Tang
arXiv:2609.38867v1 Announce Type: new
Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular in...
By Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
arXiv:2609.38285v1 Announce Type: cross
Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Add...
By Hongbo Wang, Zihan Lin, Wenkui Yang, Shiran Ge, Yuang Ai, Jie Cao, Huaibo Huang, Ran He
arXiv:2609.39145v1 Announce Type: cross
Abstract: Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how...
By Heejae Suh, Jongwook Han, Zahra Gholami, Yohan Jo