arXiv:2608. 05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding.
By S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana, Mohan Rajesh Elara
ComplicitSplat is a novel black‑box attack that leverages 3D Gaussian Splatting (3DGS) shading to create viewpoint‑specific camouflage, embedding adversarial content into scene objects that is only visible from certain angles. The method does not require access to model architecture or weights and can successfully fool a range of popular object detectors—including single‑stage, multi‑stage, and transformer‑based models—on both real‑world physical objects and synthetic scenes. This demonstrates that downstream models using 3DGS are vulnerable to adversarial manipulation.
By Matthew Hull, Haoyang Yang, Pratham Mehta, Mansi Phute, Aeree Cho, Haorang Wang, Matthew Lau, Wenke Lee, Wilian Lunardi, Martin Andreoni, Duen Horng Chau
arXiv:2603. 29418v2 Announce Type: replace-cross Abstract: Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks.
By Meiwen Ding, Song Xia, Chenqi Kong, Xudong Jiang
arXiv:2601. 14323v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed in safety-critical robotic applications, yet their security vulnerabilities remain underexplored.
By Bingxin Xu, Yuzhang Shang, Binghui Wang, Emilio Ferrara
arXiv:2608.30342v1 Announce Type: new
Abstract: 3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains lar...
By Jirong Li, Satoshi Ikehata, Shuhei Kurita, Ikuro Sato
arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.
By Jonghoon Lee, Seong Hyeon Park, Byungwoo Jeon, Minha Lee, Jinwoo Shin