arXiv Computer Vision By Ziqing Zhang, Xiao Liu, Kai Liu, Jianze Li, Weihang Zhang, Linghe Kong, Yulun Zhang

Discrete Annotation, Continuous Preference: Rethinking Supervision for Accurate and Generalizable Aesthetic Image Cropping

Read the original on arXiv Computer Vision →

The paper introduces a new approach to aesthetic image cropping by modeling human preference as a continuous, multi-peaked field rather than relying on discrete, grid‑based annotations. It presents the Continuous Preference Field (CPF) that reconstructs a dense preference landscape from sparse labels, and uses this to train a VLM‑based cropping model (CPIC) that achieves state‑of‑the‑art accuracy and strong out‑of‑domain generalization. Additionally, the authors propose CPICD, a recalibrated benchmark that corrects grid‑bound artifacts in existing datasets, providing a more reliable evaluation framework.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computation and Language
Sep 17

RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection

RankGround is a two‑stage framework for GUI grounding that uses a single Vision‑Language Model call per query. It introduces GroundRanker, a lightweight multimodal reranker that selects the most promising crop from a dense candidate set, trained with a two‑stage curriculum on ranking supervision data derived from existing grounding datasets. Experiments show RankGround outperforms strong baselines, achieving 1.4× faster inference and a 5.5% average improvement in localization accuracy over the second‑best method across all backbones and screen scales.

By Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang
arXiv Computer Vision
Aug 27

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

VGA‑BenchV2 is an expanded, human‑aligned benchmark and optimization framework that jointly evaluates video generation quality and aesthetic value. It builds on the original VGA‑Bench taxonomy, adding 52 sub‑dimensions and 1,016 curated prompts to generate over 60,000 videos from 12 mainstream models. The benchmark significantly enlarges human supervision with 36,000 task‑level annotations and introduces a hybrid evaluator (VAQA‑Net, VTag‑Net, VGQA‑Net) that aligns well with human judgments and can be used as a reward model for reinforcement‑learning fine‑tuning.

By Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin