arXiv:2607.09042v2 Announce Type: replace
Abstract: Reinforcement learning is increasingly used to fine-tune vision-language-action (VLA) models, but robot interaction is expensive and learning becom...
By Iris Xu, Sunshine Jiang, John Marangola, Pulkit Agrawal, Zhang-Wei Hong
arXiv:2609.34060v2 Announce Type: replace
Abstract: Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However...
By Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim
arXiv:2609.34457v2 Announce Type: replace
Abstract: Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, s...
By Hai Duong, Thanh Le, ThanhVu Nguyen
arXiv:2610.07774v1 Announce Type: new
Abstract: Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As m...
By Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov
arXiv:2610.08604v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address f...
By Ashley E. Bravo-Bravo, Yuchen Zhang, Haralambos Mouratidis, Ravi Shekhar, Monorama Swain
Nord-Parl-TTS is an open text‑to‑speech dataset for Finnish and Swedish created from recordings of Nordic parliamentary proceedings. It contains 900 hours of Finnish and 5 090 hours of Swedish speech, processed with an adapted Emilia pipeline and accompanied by unified evaluation sets for model development and benchmarking. The dataset aims to reduce the resource gap in TTS between high‑ and lower‑resourced languages.
By Zirui Li, Jens Edlund, Yicheng Gu, Nhan Phan, Lauri Juvela, Mikko Kurimo
The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.
By Suguru Onda, Matthew Bailey, Ryan Farrell
The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.
By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin
Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.
By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
arXiv:2610.06896v1 Announce Type: new
Abstract: Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial g...
By Ross Callaghan, Niannu Gao, Hojjat Azadbakht, Hui Zhang
arXiv:2610.07031v1 Announce Type: new
Abstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics a...
By Sitian Shen, Jiuming Liu, Mengmeng Liu, Yian Wang, Michael Ying Yang, Francesco Nex, Hao Cheng, Daniele De Martini, Ayush Tewari, Per Ola Kristensson
arXiv:2610.07689v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense p...
By Juntong Li, Lingwei Dang, Haomin Wu, Ziyan Qiu, Qingxin Xiao, Qingyao Wu
arXiv:2610.07868v1 Announce Type: new
Abstract: Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework...
By Sunchan Park, Beomkwon Cho, Kyeongbo Kong
arXiv:2610.07903v1 Announce Type: new
Abstract: Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing...
By Keliang Chen, Yaxin Hou, Hui Liu, Yuheng Jia
arXiv:2610.07925v1 Announce Type: new
Abstract: Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress,...
By Giyeol Kim, Chanho Eom
arXiv:2610.08539v1 Announce Type: new
Abstract: Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradig...
By Dongchen Si, Di Wang, Mingzhen Xu, Jing Zhang, Bo Du, Liangpei Zhang
arXiv:2610.08639v1 Announce Type: new
Abstract: As generated images become increasingly realistic, reliable forgery detection is essential for maintaining trust in visual information. However, existi...
By Jiahua Li, Zixu John, Tom Zhong, Fuping Wu, Tianhao Xu, Jianqing Zheng, Yuanhan Mo, Fei Shen
arXiv:2610.08791v1 Announce Type: new
Abstract: Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and p...
By Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin, Ziming Qin, Zheng Jiang, Wenyi Li, Calvin Xiao, Youjie Zheng, Kaisen Yang, Qinhuai Na
arXiv:2607.17669v2 Announce Type: replace
Abstract: Drone-based object detection technology has advanced rapidly, becoming increasingly sophisticated and efficient. Recently, research trends have exp...
By Hyun-Ki Jung
The paper introduces Physics‑Guided Visual Prompting (PG‑VP), a plug‑and‑play module that overlays a virtual obstacle onto the input of a frozen Vision‑Language‑Action model to guide navigation around invisible hazards such as radiation or temperature spikes. PG‑VP performs a physics‑based risk assessment to determine the avoidance direction and dynamically renders the same virtual obstacle across frames, allowing the existing navigation policy to detour without retraining. Experiments on OmniNav with R2R‑CE and RxR‑CE datasets show that PG‑VP steers the policy toward low‑risk actions in 84.9% and 83.2% of cases, while real‑world tests on a robot demonstrate significant safety improvements against thermal and radiation sources.
By Hojoon Son, Fan Zhang