Preference Optimization for Vision Language Models
Related stories
A Dive into Vision-Language Models
Vision Language Models Explained
Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
A Dataset for Dynamic Human Preferences for Vision Language Models
arXiv:2606. 07653v1 Announce Type: cross Abstract: Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models can adapt to real-time preferences for different users.
PaliGemma – Google's Cutting-Edge Open Vision Language Model
Welcome PaliGemma 2 – New vision language models by Google
SmolVLM - small yet mighty Vision Language Model
Accelerating Vision-Language Models: BridgeTower on Habana Gaudi2
Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs
arXiv:2606. 28401v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that is not grounded in the image.
P\textsuperscript{2}-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization
arXiv:2606. 03376v1 Announce Type: cross Abstract: Hallucination has recently garnered significant research attention in Large Vision-Language Models (LVLMs).
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.