Hugging Face Trending Papers

Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator

Read the original on Hugging Face Trending Papers →

Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient regions and therefore have limited sensitivity to the global relationships among the main image components.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computer Vision
4d ago

Look Closer: Patch-wise Supervision for AI-Generated Image Detection

The paper investigates patch‑wise supervision for detecting AI‑generated images, proposing a shared backbone that classifies explicit crops with individual losses and averages patch probabilities only during inference. This approach eliminates the need for handcrafted residual filtering or learned image‑level fusion modules. Experiments across single‑patch selection, multiple generator collections, and four CNN and Transformer backbones show that patch‑wise variants outperform whole‑image counterparts on the GenImage dataset, while also exploring factors such as supervision granularity, source resolution, crop size, and inference coverage.

By Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao
arXiv AI
Aug 24

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

AT‑ViT is a dual‑branch Vision Transformer that processes both raw herbarium scans and their segmentation masks through a multi‑scale, multi‑view cross‑attention fusion. It uses a mask‑guided patch weighting scheme to emphasize plant regions and suppress background artifacts, thereby encouraging plant‑centric representations. In trait classification tasks such as leaf base shape and thorns, AT‑ViT consistently outperforms baselines, improves spatial attention grounding (IoU_p +15.66 to +18.03 pp, IoU_b –27.92 to –31.02 pp), and shows greater robustness to synthetic background perturbations, surpassing ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points. whyItMatters":"The model addresses shortcut learning caused by background cues in herbarium images, leading to more accurate and interpretable plant trait recognition."

By Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti