Hugging Face Trending Papers

Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator

Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient regions and therefore have limited sensitivity to the global relationships among the main image components.

arXiv Computer Vision
4d ago

Look Closer: Patch-wise Supervision for AI-Generated Image Detection

The paper investigates patch‑wise supervision for detecting AI‑generated images, proposing a shared backbone that classifies explicit crops with individual losses and averages patch probabilities only during inference. This approach eliminates the need for handcrafted residual filtering or learned image‑level fusion modules. Experiments across single‑patch selection, multiple generator collections, and four CNN and Transformer backbones show that patch‑wise variants outperform whole‑image counterparts on the GenImage dataset, while also exploring factors such as supervision granularity, source resolution, crop size, and inference coverage.

By Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao
arXiv AI
Aug 24

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

AT‑ViT is a dual‑branch Vision Transformer that processes both raw herbarium scans and their segmentation masks through a multi‑scale, multi‑view cross‑attention fusion. It uses a mask‑guided patch weighting scheme to emphasize plant regions and suppress background artifacts, thereby encouraging plant‑centric representations. In trait classification tasks such as leaf base shape and thorns, AT‑ViT consistently outperforms baselines, improves spatial attention grounding (IoU_p +15.66 to +18.03 pp, IoU_b –27.92 to –31.02 pp), and shows greater robustness to synthetic background perturbations, surpassing ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points. whyItMatters":"The model addresses shortcut learning caused by background cues in herbarium images, leading to more accurate and interpretable plant trait recognition."

By Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti
arXiv Computer Vision
2d ago

Discrete Annotation, Continuous Preference: Rethinking Supervision for Accurate and Generalizable Aesthetic Image Cropping

The paper introduces a new approach to aesthetic image cropping by modeling human preference as a continuous, multi-peaked field rather than relying on discrete, grid‑based annotations. It presents the Continuous Preference Field (CPF) that reconstructs a dense preference landscape from sparse labels, and uses this to train a VLM‑based cropping model (CPIC) that achieves state‑of‑the‑art accuracy and strong out‑of‑domain generalization. Additionally, the authors propose CPICD, a recalibrated benchmark that corrects grid‑bound artifacts in existing datasets, providing a more reliable evaluation framework.

By Ziqing Zhang, Xiao Liu, Kai Liu, Jianze Li, Weihang Zhang, Linghe Kong, Yulun Zhang
arXiv AI
Aug 11

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

arXiv:2512. 05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes.

By Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang