arXiv:2608. 01301v3 Announce Type: replace-cross Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize proxies for information transfer, structure, or source similarity.
By Haoran Liu, Mingzhe Liu, Peng Li, Guibin Zan
The paper introduces the Learned Perceptual Image Fusion Measure (LPIFM), a model trained on dense human pairwise comparisons to assess infrared-visible image fusion. LPIFM jointly processes both source images and fused candidates using a shared hierarchical encoder, triadic interaction, and a tie-aware objective, achieving high agreement with human judgments and outperforming 19 conventional metrics. The authors release a large comparison corpus, model weights, and code, demonstrating LPIFM’s rapid adaptability to new fusion-evaluation protocols.
By Haoran Liu, Mingzhe Liu, Peng Li, Guibin Zan
VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.
By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
PreResQ‑R1 introduces a Preference‑Response Disentangled Reinforcement Learning framework for Visual Quality Assessment that jointly optimizes absolute score regression and relative ranking consistency. It employs a dual‑branch reward system—modeling intra‑sample response coherence and inter‑sample preference alignment—trained with Group Relative Policy Optimization. The method extends to video quality assessment via a global‑temporal and local‑spatial data flow strategy, achieving state‑of‑the‑art results on 10 IQA and 5 VQA benchmarks with only 6K images and 28K videos, and provides human‑aligned reasoning traces.
By Zehui Feng, Weichuan Wang, Xiaohan Chen, Ting Han
arXiv:2609.25716v1 Announce Type: new
Abstract: Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn ho...
By Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.
By Amit Jadhav, Shaurya Beriwala, Beomjin Kim