arXiv:2608. 07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred.
By Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li, Yinyin Gong, Yipo Huang, Jiliang Zhao
arXiv:2607. 06915v1 Announce Type: cross Abstract: Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI.
By Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin
The paper investigates patch‑wise supervision for detecting AI‑generated images, proposing a shared backbone that classifies explicit crops with individual losses and averages patch probabilities only during inference. This approach eliminates the need for handcrafted residual filtering or learned image‑level fusion modules. Experiments across single‑patch selection, multiple generator collections, and four CNN and Transformer backbones show that patch‑wise variants outperform whole‑image counterparts on the GenImage dataset, while also exploring factors such as supervision granularity, source resolution, crop size, and inference coverage.
By Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao
arXiv:2603.27519v4 Announce Type: replace
Abstract: Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, orga...
By Shuai Xiang, James Burridge, Shouyang Liu, Hao Lu, Tokihiro Fukatsu, Yinqiang Zheng, Wei Guo
AT‑ViT is a dual‑branch Vision Transformer that processes both raw herbarium scans and their segmentation masks through a multi‑scale, multi‑view cross‑attention fusion. It uses a mask‑guided patch weighting scheme to emphasize plant regions and suppress background artifacts, thereby encouraging plant‑centric representations. In trait classification tasks such as leaf base shape and thorns, AT‑ViT consistently outperforms baselines, improves spatial attention grounding (IoU_p +15.66 to +18.03 pp, IoU_b –27.92 to –31.02 pp), and shows greater robustness to synthetic background perturbations, surpassing ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points.
whyItMatters":"The model addresses shortcut learning caused by background cues in herbarium images, leading to more accurate and interpretable plant trait recognition."
By Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti
arXiv:2609.38165v1 Announce Type: cross
Abstract: The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However,...
By Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini
The paper introduces a new approach to aesthetic image cropping by modeling human preference as a continuous, multi-peaked field rather than relying on discrete, grid‑based annotations. It presents the Continuous Preference Field (CPF) that reconstructs a dense preference landscape from sparse labels, and uses this to train a VLM‑based cropping model (CPIC) that achieves state‑of‑the‑art accuracy and strong out‑of‑domain generalization. Additionally, the authors propose CPICD, a recalibrated benchmark that corrects grid‑bound artifacts in existing datasets, providing a more reliable evaluation framework.
By Ziqing Zhang, Xiao Liu, Kai Liu, Jianze Li, Weihang Zhang, Linghe Kong, Yulun Zhang
arXiv:2505.16157v3 Announce Type: replace
Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
By Yuang Ai
arXiv:2608.22996v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution s...
By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang
arXiv:2512. 05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes.
By Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang
arXiv:2608.03423v2 Announce Type: replace
Abstract: Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstru...
By Zhihua Xu, Runyu Zhu, Rongjun Qin
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in light...