Low-light image enhancement (LLIE) methods involve tunable parameters that are typically fixed, often leading to performance degradation when applied across scenes. Manually selecting the best configuration, however, can be time-consuming and not always practical.
LLaVA‑Assessor is a unified large multi‑modal model (LMM) designed for visual quality assessment, combining image and video inputs. It introduces a two‑task framework—quality interpretation and quality scoring—supported by an adaptive architecture, a rigorous human‑annotated dataset, and a machine‑synthesized data expansion pipeline. The model employs a prompt‑disentanglement strategy to stabilize multi‑task training and achieves strong performance across 11 quality scoring test sets and 4 interpretation benchmarks.
By Ziheng Jia, Zicheng Zhang, Jiaying Qian, Guangtao Zhai, Xiongkuo Min
Q‑SiT is a unified framework that trains large multimodal models to perform both image quality scoring and interpreting simultaneously. By converting standard IQA datasets into question‑answer pairs and adding human‑annotated interpreting data, the model learns to quantify overall quality and describe perceived attributes. An efficient balance strategy optimizes data mix ratios on lightweight models before scaling to full‑size LMMs, reducing computational cost while improving cross‑task knowledge transfer.
By Zicheng Zhang, Haoning Wu, Ziheng Jia, Weisi Lin, Guangtao Zhai
arXiv:2606. 16082v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA).
By Guanyi Qin, Junjie Zhang, Chunming He, Yibing Fu, Jie Liang, Tianhe Wu, Lei Zhang
arXiv:2609.14495v1 Announce Type: new
Abstract: Image colorization is an inherently ill-posed task, since a single grayscale image may correspond to multiple plausible colorized results. Consequently...
By Yunkai Zhuang, Qihang Yan, Zicheng Zhang, Guangtao Zhai
arXiv:2512. 05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes.
By Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang
arXiv:2607. 03013v1 Announce Type: cross Abstract: Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to sensor limitations and imaging pipelines, which degrades visual quality and affects downstream vision tasks.
By Wanshu Fan, Xiangyu Li, Cong Wang, Kin-man Lam, Xin Yang, Haiyan Zhang, Dongsheng Zhou
IDM-Net is a lightweight Illumination-Decoupled Modulation Network designed for low-light image enhancement. It uses a dual-encoder architecture: a structure encoder extracts multi-scale appearance features from the RGB image, while a lightweight illumination encoder learns illumination priors from the decoupled luminance channel. An Illumination-Guided Modulation module injects these priors into the decoder via spatially adaptive affine modulation, and a Feature Refinement Block progressively suppresses artifacts and recovers fine details, achieving competitive performance on standard benchmarks with good balance between quality and efficiency.
By Cheng-Yen Hsiao, Jing-Ming Guo
PreResQ‑R1 introduces a Preference‑Response Disentangled Reinforcement Learning framework for Visual Quality Assessment that jointly optimizes absolute score regression and relative ranking consistency. It employs a dual‑branch reward system—modeling intra‑sample response coherence and inter‑sample preference alignment—trained with Group Relative Policy Optimization. The method extends to video quality assessment via a global‑temporal and local‑spatial data flow strategy, achieving state‑of‑the‑art results on 10 IQA and 5 VQA benchmarks with only 6K images and 28K videos, and provides human‑aligned reasoning traces.
By Zehui Feng, Weichuan Wang, Xiaohan Chen, Ting Han
The paper presents MM‑IQA, a lightweight no‑reference image quality assessment framework designed for UAV imaging. It fuses interpretable metrics—blur, edge structure, low‑resolution artifacts, exposure imbalance, noise, haze, and frequency content—to output a single quality score between 0 and 100. Evaluated on five benchmark datasets, MM‑IQA achieved SRCC values from 0.647 to 0.830 and runs in about 1.97 s per image with modest memory usage.
By Koffi Titus Sergio Aglin, Anthony K. Muchiri, Celestin Nkundineza
Consist‑Retinex introduces a one‑step noise‑emphasized consistency training framework for Retinex‑based low‑light image enhancement. It first decomposes images into reflectance and illumination maps using a Retinex Transformer Decomposition Network, then trains two conditional consistency models with a dual objective that blends trajectory consistency and ground‑truth alignment. The method employs adaptive noise‑emphasized fixed‑point sampling to focus supervision near the inference endpoint, achieving state‑of‑the‑art VE‑LOL‑L scores on paired and unpaired low‑light benchmarks while reducing sampling and training costs.
By Jian Xu, Wei Chen, Shigui Li, Delu Zeng, John Paisley, Qibin Zhao
The paper introduces a new task called Quality Anomaly Perception for UGC Image Enhancement (UEAP) and presents the first benchmark dataset, UEAP-4k, featuring fine‑grained annotations of anomaly categories, locations, and severity levels in real‑world user‑generated content. It proposes the Difference‑Fusion Anomaly Perception Method (DFAP‑UGC), which fuses explicit differences between enhanced images and their references using dense spatial querying, regional verification, and quality‑aware ranking to robustly identify localized anomalies. A Locality‑Aware Dynamic Task Prioritization (LADTP) training strategy is also introduced to enable efficient end‑to‑end learning without multi‑stage overhead, and experiments demonstrate that DFAP‑UGC outperforms adapted classical baselines.
By Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi, Tingting Jiang