arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri
arXiv:2603. 29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment.
By Qiucheng Yu, Ruijie Xu, Mingang Chen Jianfeng Dong, Xin Tan
arXiv:2607. 00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories.
By Siddhant Panpatil, Arth Singh, Mijin Koo, Chaeyun Kim, Haon Park, Dasol Choi
arXiv:2609.13225v1 Announce Type: cross
Abstract: Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fail...
By Sarthak Sattigeri
arXiv:2509. 25533v2 Announce Type: replace-cross Abstract: As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns has become increasingly important.
By Ravikumar Balakrishnan, Mansi Phute
Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\textit{InternVL}, \textit{Qwen-VL}, and \textit{SmolVLM}) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tuning against a stratified random baseline.
arXiv:2608.23313v1 Announce Type: new
Abstract: Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view c...
By Xuetong Li, Gaofeng Liu
arXiv:2608.21832v1 Announce Type: new
Abstract: Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether mo...
By Md Abrar Jahin, Md Rizwan Parvez
Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.
arXiv:2609.00868v1 Announce Type: cross
Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...
By Genpei Zhang
AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon evaluation framework for vision‑language models that tests hidden‑target identification across galleries of visually similar images. The benchmark requires a model to ask a sequence of attribute‑focused Yes/No questions, receiving Yes/No/Unsure feedback and penalizing invalid actions with Skip, thereby stressing question selection under uncertainty, constraint tracking, and fine‑grained discrimination. Using the Guess My Preferred Dress task, the study shows that final‑answer accuracy alone overstates performance, as models may guess correctly without verified evidence, waste turns, or violate the protocol, highlighting the need for combined visual discrimination, informative questioning, and robust protocol adherence.
By Min Wang, Ata Mahjoubfar
arXiv:2607. 22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting.
By M M Asif Ferdous