YOLO with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection with trustworthy multimodal AI in computer vision perception
Read the original on arXiv AI →The paper introduces a Kolmogorov-Arnold network as an interpretable post‑hoc surrogate to assess the trustworthiness of YOLOv10 object detections, using seven geometric and semantic features. Its additive spline structure allows direct visualization of each feature’s influence, revealing when confidence scores are reliable or unreliable. A bootstrapped BLIP language‑image model generates scene captions, providing a lightweight multimodal interface that does not compromise interpretability. Experiments on COCO and University of Bath campus images show the framework accurately flags low‑trust predictions under blur, occlusion, or low texture, offering actionable insights for risk mitigation.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.