UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
Read the original on arXiv AI →UltraG-Bench is a large‑scale, multi‑task benchmark designed to evaluate pixel‑level evidence grounding in ultrasound images. It comprises 40 public segmentation datasets covering 13 anatomical categories and includes three progressive tasks—instruction‑guided segmentation, evidence‑grounded VQA, and evidence‑grounded report generation—with a total of 736,726 annotations. Evaluation of 14 state‑of‑the‑art models shows a significant gap between semantic understanding and fine‑grained pixel‑level localization, and the authors propose UltraG‑Agent, which combines a vision‑language model with the ultrasound‑specific segmentation model UltraSAM to improve both semantic prediction and visual grounding.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.