SABRE: Scalable and Automated Benchmarking of VLMs under Stress
arXiv:2608. 07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify.
arXiv:2606. 16082v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA).
arXiv:2608. 07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify.
LLaVA‑Assessor is a unified large multi‑modal model (LMM) designed for visual quality assessment, combining image and video inputs. It introduces a two‑task framework—quality interpretation and quality scoring—supported by an adaptive architecture, a rigorous human‑annotated dataset, and a machine‑synthesized data expansion pipeline. The model employs a prompt‑disentanglement strategy to stabilize multi‑task training and achieves strong performance across 11 quality scoring test sets and 4 interpretation benchmarks.
arXiv:2601. 02918v4 Announce Type: replace Abstract: Image Quality Assessment (IQA) is a long-standing problem in computer vision.
arXiv:2608. 09111v1 Announce Type: new Abstract: AI video generation has advanced rapidly and entered widespread commercial use.
arXiv:2607. 12375v1 Announce Type: cross Abstract: Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability.
The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following.
arXiv:2609.37576v1 Announce Type: new Abstract: With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to captur...
VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.
arXiv:2605.26380v2 Announce Type: replace-cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
arXiv:2605.21244v2 Announce Type: replace Abstract: Super-Resolution (SR) has advanced rapidly in recent years, with diffusion-based models achieving unprecedented fidelity at the cost of introducing...
arXiv:2607. 10826v1 Announce Type: cross Abstract: Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow.