Hugging Face Blog Jun 24, 2024 Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
Hugging Face Blog Jun 3, 2025 SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data
Hugging Face Blog Apr 15, 2024 Introducing Idefics2: A Powerful 8B Vision-Language Model for the community
arXiv AI Jul 28 From Pixels to Prompts: Vision-Language Models arXiv:2605. 07544v3 Announce Type: replace Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. By Khang Nhat Hoang Vo