Introducing IDEFICS: An Open Reproduction of State-of-the-art Visual Langage Model
Related stories
A Dive into Vision-Language Models
Vision Language Models Explained
Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
Vision Language Models (Better, faster, stronger)
A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
Accelerating Vision-Language Models: BridgeTower on Habana Gaudi2
Vision-Language Models are Fragile Multilingual Associators
arXiv:2608. 12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes.
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
The paper reports a submission to the SHROOM-Visions shared task, aiming to detect and classify hallucinated character spans in vision‑language model outputs across four languages. The authors use multiple fine‑tuned vision‑language models as independent annotators, combine their predictions via character‑level majority voting, and also investigate activation probes. Their method achieved first place in three of the four languages and consistently ranked on the podium for all languages and metrics, with analysis showing that model disagreement mirrors human annotator disagreement.