Explainable deep learning improves human mental models of self-driving cars
arXiv:2411. 18714v3 Announce Type: replace-cross Abstract: Self-driving cars increasingly rely on deep neural networks to achieve human-like driving.
arXiv:2411. 18714v3 Announce Type: replace-cross Abstract: Self-driving cars increasingly rely on deep neural networks to achieve human-like driving.
arXiv:2509.22404v2 Announce Type: replace Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
arXiv:2607. 24730v1 Announce Type: cross Abstract: Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust.
arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.
Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from severe performance degradation when deployed on novel scenes, unseen categories, or visually confusing backgrounds. Moreover, existing unified paradigms primarily rely on intra-image specific prompts, lacking flexible task routing to adapt to multi-intent operational workflows.
arXiv:2608. 12299v1 Announce Type: cross Abstract: Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence.
arXiv:2411. 05698v3 Announce Type: replace-cross Abstract: Convolutional Neural Networks (CNNs) have shown remarkable performance in image classification.
arXiv:2606. 07180v1 Announce Type: cross Abstract: The growing demand for transparency in automated decision-making has propelled eXplainable Artificial Intelligence (XAI) to the forefront of machine learning research.
arXiv:2605. 16651v2 Announce Type: replace-cross Abstract: Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model decisions require human oversight.
Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations.
arXiv:2512. 21414v2 Announce Type: replace-cross Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools.
The paper introduces a Kolmogorov-Arnold network as an interpretable post‑hoc surrogate to assess the trustworthiness of YOLOv10 object detections, using seven geometric and semantic features. Its additive spline structure allows direct visualization of each feature’s influence, revealing when confidence scores are reliable or unreliable. A bootstrapped BLIP language‑image model generates scene captions, providing a lightweight multimodal interface that does not compromise interpretability. Experiments on COCO and University of Bath campus images show the framework accurately flags low‑trust predictions under blur, occlusion, or low texture, offering actionable insights for risk mitigation.