LLaVA‑Assessor is a unified large multi‑modal model (LMM) designed for visual quality assessment, combining image and video inputs. It introduces a two‑task framework—quality interpretation and quality scoring—supported by an adaptive architecture, a rigorous human‑annotated dataset, and a machine‑synthesized data expansion pipeline. The model employs a prompt‑disentanglement strategy to stabilize multi‑task training and achieves strong performance across 11 quality scoring test sets and 4 interpretation benchmarks.
By Ziheng Jia, Zicheng Zhang, Jiaying Qian, Guangtao Zhai, Xiongkuo Min
arXiv:2608.24372v1 Announce Type: new
Abstract: AI-generated image quality assessment (AIGIQA) requires jointly reasoning about perceptual fidelity and prompt alignment, two quality dimensions that a...
By Baoliang Chen, Qing Lin, Sijie Mai
PreResQ‑R1 introduces a Preference‑Response Disentangled Reinforcement Learning framework for Visual Quality Assessment that jointly optimizes absolute score regression and relative ranking consistency. It employs a dual‑branch reward system—modeling intra‑sample response coherence and inter‑sample preference alignment—trained with Group Relative Policy Optimization. The method extends to video quality assessment via a global‑temporal and local‑spatial data flow strategy, achieving state‑of‑the‑art results on 10 IQA and 5 VQA benchmarks with only 6K images and 28K videos, and provides human‑aligned reasoning traces.
By Zehui Feng, Weichuan Wang, Xiaohan Chen, Ting Han
arXiv:2607. 18237v1 Announce Type: cross Abstract: Human visual similarity judgments are context-dependent.
By Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi
arXiv:2609.14495v1 Announce Type: new
Abstract: Image colorization is an inherently ill-posed task, since a single grayscale image may correspond to multiple plausible colorized results. Consequently...
By Yunkai Zhuang, Qihang Yan, Zicheng Zhang, Guangtao Zhai
arXiv:2607. 12375v1 Announce Type: cross Abstract: Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability.
By Jinjian Wu, Jiaqi Tang, Wei Wei, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen
arXiv:2603.08064v3 Announce Type: replace
Abstract: Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are e...
By Zexi Jia, Pengcheng Luo, Yijia Zhong, Jinchao Zhang, Jie Zhou
arXiv:2506.02015v4 Announce Type: replace
Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still strug...
By Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi, Sungwoong Kim
arXiv:2606. 16082v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA).
By Guanyi Qin, Junjie Zhang, Chunming He, Yibing Fu, Jie Liang, Tianhe Wu, Lei Zhang
VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.
By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
arXiv:2609.22942v1 Announce Type: new
Abstract: Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plau...
By Zhenchen Tang, Bo Peng, Zichuan Wang, Songlin Yang, Leilei Cao, Fengjie Zhu, Jing Dong