arXiv Computer Vision

PiPS: Post-Hoc Prototypical Explanations for Interpretable Semantic Segmentation

arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
Hugging Face Trending Papers
Aug 4

UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution

Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from severe performance degradation when deployed on novel scenes, unseen categories, or visually confusing backgrounds. Moreover, existing unified paradigms primarily rely on intra-image specific prompts, lacking flexible task routing to adapt to multi-intent operational workflows.

Hugging Face Trending Papers
Jun 23

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations.

arXiv AI
Sep 7

YOLO with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection with trustworthy multimodal AI in computer vision perception

The paper introduces a Kolmogorov-Arnold network as an interpretable post‑hoc surrogate to assess the trustworthiness of YOLOv10 object detections, using seven geometric and semantic features. Its additive spline structure allows direct visualization of each feature’s influence, revealing when confidence scores are reliable or unreliable. A bootstrapped BLIP language‑image model generates scene captions, providing a lightweight multimodal interface that does not compromise interpretability. Experiments on COCO and University of Bath campus images show the framework accurately flags low‑trust predictions under blur, occlusion, or low texture, offering actionable insights for risk mitigation.

By Marios Impraimakis, Daniel Vazquez, Feiyu Zhou