arXiv Computer Vision

An Evolutionary Agentic Approach for Open-ended Image Quality Perception

arXiv AI
Sep 15

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

NoteVQA is a new benchmark that collects 252 real‑life visual questions from the Chinese image‑sharing platform Xiaohongshu, covering 12 topics and 7 user intents. Each question is paired with a concise expert reference and a human‑audited interleaved answer that blends text and visual evidence. The study evaluates VLMs on short‑answer correctness and interleaved answer quality using a new AgenticInterleave framework and a 12‑dimensional IVR‑12 rubric, finding that even state‑of‑the‑art models achieve only about 53% accuracy and lag behind human references in content quality.

By Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen, Yahui Liu, Yao Hu, Chuan Mu
arXiv AI
Jun 26

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

arXiv:2606. 26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images.

By Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu
arXiv Computer Vision
Aug 26

EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI

EgoErrorVQA introduces a new egocentric visual question answering task that evaluates visual agents’ ability to detect procedural errors in everyday activities. The paper presents an evaluator agent built on the Agent2Agent protocol and shows that current models struggle with procedural error recognition. It also proposes Ego-ADR, an Adaptive Decoupled Reasoning framework that improves performance on the task, achieving state‑of‑the‑art results.

By Junlong Li, Junxi Li, Jianjun Gao, Chen Cai, Lap-Pui Chau, Yi Wang
arXiv Computer Vision
Aug 27

PreResQ-R1: Response-Preference Disentangled Ranking-and-Scoring Reinforcement Optimization for Robust Visual Quality Assessment

PreResQ‑R1 introduces a Preference‑Response Disentangled Reinforcement Learning framework for Visual Quality Assessment that jointly optimizes absolute score regression and relative ranking consistency. It employs a dual‑branch reward system—modeling intra‑sample response coherence and inter‑sample preference alignment—trained with Group Relative Policy Optimization. The method extends to video quality assessment via a global‑temporal and local‑spatial data flow strategy, achieving state‑of‑the‑art results on 10 IQA and 5 VQA benchmarks with only 6K images and 28K videos, and provides human‑aligned reasoning traces.

By Zehui Feng, Weichuan Wang, Xiaohan Chen, Ting Han
arXiv Computer Vision
Sep 15

Q-SiT: Teaching LMMs for Image Quality Scoring and Interpreting

Q‑SiT is a unified framework that trains large multimodal models to perform both image quality scoring and interpreting simultaneously. By converting standard IQA datasets into question‑answer pairs and adding human‑annotated interpreting data, the model learns to quantify overall quality and describe perceived attributes. An efficient balance strategy optimizes data mix ratios on lightweight models before scaling to full‑size LMMs, reducing computational cost while improving cross‑task knowledge transfer.

By Zicheng Zhang, Haoning Wu, Ziheng Jia, Weisi Lin, Guangtao Zhai