arXiv:2608. 07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred.
By Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li, Yinyin Gong, Yipo Huang, Jiliang Zhao
RankGround is a two‑stage framework for GUI grounding that uses a single Vision‑Language Model call per query. It introduces GroundRanker, a lightweight multimodal reranker that selects the most promising crop from a dense candidate set, trained with a two‑stage curriculum on ranking supervision data derived from existing grounding datasets. Experiments show RankGround outperforms strong baselines, achieving 1.4× faster inference and a 5.5% average improvement in localization accuracy over the second‑best method across all backbones and screen scales.
By Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang
We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-Be...
VGA‑BenchV2 is an expanded, human‑aligned benchmark and optimization framework that jointly evaluates video generation quality and aesthetic value. It builds on the original VGA‑Bench taxonomy, adding 52 sub‑dimensions and 1,016 curated prompts to generate over 60,000 videos from 12 mainstream models. The benchmark significantly enlarges human supervision with 36,000 task‑level annotations and introduces a hybrid evaluator (VAQA‑Net, VTag‑Net, VGQA‑Net) that aligns well with human judgments and can be used as a reward model for reinforcement‑learning fine‑tuning.
By Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin
arXiv:2603.27519v4 Announce Type: replace
Abstract: Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, orga...
By Shuai Xiang, James Burridge, Shouyang Liu, Hao Lu, Tokihiro Fukatsu, Yinqiang Zheng, Wei Guo
arXiv:2606. 08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste.
By Harini SI, Somesh Singh, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah