Hugging Face Blog
Jun 24, 2024
The paper reports a submission to the SHROOM-Visions shared task, aiming to detect and classify hallucinated character spans in vision‑language model outputs across four languages. The authors use multiple fine‑tuned vision‑language models as independent annotators, combine their predictions via character‑level majority voting, and also investigate activation probes. Their method achieved first place in three of the four languages and consistently ranked on the podium for all languages and metrics, with analysis showing that model disagreement mirrors human annotator disagreement.
arXiv:2608. 12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes.