The paper presents VLANeXt, a Vision‑Language‑Action (VLA) model built from a unified framework that systematically analyzes design choices across foundational components, perception essentials, and action modeling. By dissecting these dimensions, the authors identify 12 key findings that form a practical recipe for strong VLA models, and demonstrate that VLANeXt outperforms state‑of‑the‑art methods on LIBERO and LIBERO‑plus benchmarks while also excelling in real‑world experiments. The study further extends VLANeXt into a family of variants—scaled models, latent‑action pretraining, latent predictive representation learning, and world action modeling—showing that the core recipe remains effective across different scales and emerging paradigms.
By Xiao-Ming Wu, Kang Liao, Yihang Luo, Bin Fan, Jian-Jian Jiang, Runze Yang, Zhonghua Wu, Wei-Shi Zheng, Chen Change Loy
The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.
By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
The paper presents a preprocessor that recovers and decodes encoded content in vision‑language models to close the decode gap that allows harmful requests to bypass safety classifiers. Evaluated against eleven encoding attacks, the preprocessor raises block rates from 0 % to 67‑90 % but also increases benign over‑refusal, and no configuration achieves an ensemble attack‑success rate below 40 % while keeping benign over‑refusal under 70 %. The study shows that closing one encoding channel merely relocates success rather than eliminating it, highlighting the limits of recovery‑based defenses.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng, Haowen Xu, Xiangchen Guan, Yang Chen, Zijian Xiao, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita
The paper demonstrates that vision‑language models’ refusal behavior is heavily influenced by whether an image is attached to a request, even when the image is blank or unreadable. Attaching such an image shifts refusal scores by large margins for borderline‑benign prompts while leaving genuinely neutral instructions largely unchanged. This effect varies with image properties, persists across checkpoints, and is not mitigated by explicit instructions to ignore the image.
By Haoyu Zhang, Yi Feng, Hanwen Liu, Shibo Zheng, Zhuoxi Wang, Yang Chen, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2609.30629v1 Announce Type: cross
Abstract: Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication const...
By Rajat Bhattacharjya, Minwoo Kim, Arnab Sarkar, Tamoghno Das, Sing-Yao Wu, Eli Bozorgzadeh, Marco Levorato, Nikil Dutt
arXiv:2609.31356v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as ideali...
By Sumanth Udupa, Mehrtash Harandi, Yadan Luo, Mahsa Baktashmotlagh
The paper introduces Muslim, an Arabic voice AI platform that delivers grounded Islamic knowledge to users in real time. It combines a NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, and a self-hosted TTS system, supported by a deterministic multi-source retrieval layer across six Model Context Protocol servers. The authors release fine‑tuned Arabic Islamic model artifacts, describe an account‑based metering layer that prevents abuse, and present a three‑layer observability stack that monitors GPU‑bound agent hosts, reporting latency, accuracy, and engineering trade‑offs for production deployment.
By Yahya Mohamed Elnawasany
Audio large language models (Audio LLMs) often fail to transcribe English‑Mandarin code‑switching speech, exhibiting language omission, translation‑instead‑of‑transcription, and hallucination. By applying Direct Preference Optimization (DPO) with 100 K preference pairs, the models learn to preserve mixed‑language content rather than translate, leading to significant reductions in mixed‑error rates (up to 89.6% in‑distribution). The study demonstrates that DPO can effectively align multilingual Audio LLMs for accurate code‑switching transcription.
By Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun, Ai Ti Aw
The paper investigates prompt injection attacks on 14 open‑source and 3 closed‑source large language models (LLMs), introducing a new metric called Attack Success Probability (ASP) that accounts for uncertainty in model responses. It demonstrates that a simple hypnotism attack can trigger objectionable behavior in models such as StableLM2, Mistral, Openchat, and Vicuna, achieving roughly 90% ASP. The study highlights that moderately well‑known LLMs are particularly vulnerable, underscoring the importance of public awareness and effective mitigation strategies.
By Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen
ProCAP introduces a probabilistic cross-attentive prompt learning framework for vision-language models like CLIP, enabling improved cross-modal interaction without updating the backbone. It jointly learns visual and textual prompt tokens, linking them via stacked bidirectional multi-head cross-attention to refine each branch across prompt depth. The method incorporates Gaussian parameterization of prompt tokens, lightweight KL and L2 regularization, and a compact symmetric InfoNCE head to align image features with class-level text representations, achieving strong few-shot base-to-novel performance and competitive transfer results across multiple datasets and benchmarks.
By Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
EviDETR is a new framework for joint video moment retrieval and highlight detection that preserves query‑relevant temporal evidence throughout its pipeline. It introduces three key components: Semantic‑aware Feature Reweighting (SFR) to enhance clip representations, a Temporal Top‑2 Mixture‑of‑Experts (TTop2MoE) decoder for query‑adaptive refinement, and MR‑to‑HD (MR2HD) fusion to transfer retrieval evidence to highlight prediction. Using CLIP+SlowFast features, EviDETR achieves state‑of‑the‑art performance on QVHighlights and demonstrates strong cross‑dataset transferability on TACoS and Charades‑STA.
By Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang
ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.
By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li
TRACKGRAPH is an online open‑vocabulary 3D mapping system that tracks 2D masks in the image stream before fusing them into a class‑agnostic 3D segment layer within a hierarchical scene graph. It uses FastSAM and CLIP for sparse keyframes, DINOv3 for dense mask propagation, and compact multi‑view CLIP embeddings for open‑vocabulary retrieval. The method outperforms state‑of‑the‑art mapping techniques on Replica, ScanNet++, and HM3D, achieving higher synonym frequency, faster processing, and lower GPU memory usage, and has been deployed on quadruped robots and drones at real‑time rates.
By Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl
The paper introduces CytoCRF, a conditional random field framework tailored for cytology images. It adapts pairwise terms to focus on chromatin and cytology-specific staining and enriches neighborhood information by combining multiple backbone models. Across ten cytology datasets, CytoCRF surpasses existing CRF methods at all annotation budgets, achieving up to +13.6 percentage points over the best baseline and +33.7 over zero‑shot performance with only 50 annotations.
By Manon Dausort, Tiffanie Godelaine, Karim El Khoury, Maxime Zanella, Christophe De Vleeschouwer, Beno\^it Macq
The paper introduces SlideTIM, a transductive few‑shot classification method tailored for whole‑slide images (WSIs). SlideTIM extends the LC‑TIM approach by adding a spatial‑latent regularizer and a class‑distribution prior, ensuring that spatially and semantically similar patches receive consistent predictions and that predicted class proportions are calibrated. Experiments on four histology datasets show that SlideTIM outperforms existing TIM variants, boosting macro‑F1 scores by up to 8.1 percentage points over the best baseline and 19.4 percentage points over zero‑shot predictions at one shot.
By Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq, Christophe De Vleeschouwer
The paper evaluates pseudo‑labeling to adapt pretrained ASR models (Whisper and Qwen3‑ASR) for noisy Broadcast Police Communication (BPC) from Baltimore and Chicago. It finds that internal confidence metrics cannot reliably separate high‑ and low‑quality pseudo‑labels, and proposes an external LLM‑as‑a‑judge filtering approach that more aggressively removes implausible transcripts, reducing WER. Additionally, a cross‑model pseudo‑labeling strategy is introduced, where one model is fine‑tuned with pseudo‑labels from the other, showing promise for future work.
By Kaavya Chaparala, Su Huang, Stephen L. Miller, Rhiannon N. Miller, Anjalie Field
SEA-CLIP-Tiny is a compact multilingual text‑vision embedding model designed for Southeast Asian languages, containing fewer than 50 million parameters. It adapts a CLIP‑KD framework with region‑specific data curation and multilingual teacher guidance. Across seven languages, it outperforms other student models, achieving R@1 = 12.9%, R@5 = 31.5%, and R@10 = 42.2%, and surpasses MobileCLIP2 by 12.1 points in R@10 while using 38.4% fewer parameters and lower CPU latency.
By Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda, Peerat Limkonchotiwat
Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on R...
The paper introduces a personalized Korean visual speech recognition system that uses a video-only Conformer model initialized from English-trained weights, achieving a character error rate (CER) of 9.95–12.19% on the OLKAVS nine-camera corpus and 19.00–21.52% on unseen wording. Individual speaker CER varies widely (1.0–52.2%), with seen wording reducing errors by 7.0–9.0 points and professional or spontaneous speech increasing errors by 8.5–12.7 points. A low‑rank adapter, comprising only 4.6% of the model parameters and trained on 4–29 minutes of a user’s frontal video, reduces high‑error speakers’ CER by 2.13–3.58 points, transfers across all cameras without loss, and retains 85% of full fine‑tuning benefits at 12% of its cost; cameras above the mouth plane add a constant offset of about six CER points that can be mitigated by training on all views.
By Se Un Park, Hakjun Kim, Taehoon Roh, Junyoung Park
The paper introduces a dual‑input, multi‑task learning framework that jointly segments and classifies bone tumors by applying bidirectional cross‑modal attention between a lesion crop and the full radiograph. Using a YOLO‑based detector and a dual‑stream DenseNet121 architecture, the model fuses fine‑grained lesion detail with global anatomical context through a novel cross‑modal attention fusion strategy and hierarchical multi‑scale feature fusion. On the multi‑institutional Bone Tumor X‑ray Radiograph Dataset, the approach outperforms single‑input baselines, achieving a Dice coefficient of 0.896 and a macro‑averaged F1‑score of 0.928, with an AUC of 0.999 for malignant osteosarcoma.
By S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul