arXiv AI

Multi-Scale ViT Inference with Habitat-Fit Priors and kNN Retrieval for Multi-Species Plant Identification

arXiv:2607. 14509v1 Announce Type: cross Abstract: This paper describes DS@GT ARC's third-place solution to the PlantCLEF 2026 challenge on multi-species plant identification in vegetation quadrat images, where systems must predict every species present in high-resolution (~3000 x 3000 pixel) plot photographs while training only on single-label images of individual plants.

arXiv AI
Sep 25

AgriCountDINO: Parameter-Efficient Exemplar-Guided Counting and Localization in Agriculture

AgriCountDINO is a parameter‑efficient, exemplar‑guided framework that jointly counts and localizes plants and their organs by conditioning frozen multiscale DINOv3 features on exemplar appearance and size, then decoding them into target points. It introduces missed‑object recovery and exemplar‑adaptive point NMS to improve detection accuracy. With only 8.4 M trainable parameters, it achieves a three‑shot MAE of 11.92 on the TPC‑268 benchmark and a zero‑shot MAE of 14.25 on unseen generic object categories in FSC‑147, outperforming previous methods without target‑domain training.

By Shengjie Guo, Xin Li, Borjana Arsova, Hanno Scharr, Silvio Salvi
arXiv AI
Aug 24

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

AT‑ViT is a dual‑branch Vision Transformer that processes both raw herbarium scans and their segmentation masks through a multi‑scale, multi‑view cross‑attention fusion. It uses a mask‑guided patch weighting scheme to emphasize plant regions and suppress background artifacts, thereby encouraging plant‑centric representations. In trait classification tasks such as leaf base shape and thorns, AT‑ViT consistently outperforms baselines, improves spatial attention grounding (IoU_p +15.66 to +18.03 pp, IoU_b –27.92 to –31.02 pp), and shows greater robustness to synthetic background perturbations, surpassing ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points. whyItMatters":"The model addresses shortcut learning caused by background cues in herbarium images, leading to more accurate and interpretable plant trait recognition."

By Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti
arXiv Machine Learning
Aug 27

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

CropCop is a closed‑set plant‑health recognition system covering 120 operational classes, built from a rigorously audited dataset of 109,107 images after removing 3,233 duplicate relationships. The model, based on a fine‑tuned DINOv3 ConvNeXt‑Tiny, achieves 98.51% accuracy and 96.87% macro‑F1 on a locked internal test, while a quantised MobileNetV4 variant reaches 98.46% accuracy and 96.23% macro‑F1 in a 22.60 MiB runtime artifact. Validation‑only post‑training quantisation and a compact ExecuTorch/XNNPACK PTE ensure high fidelity between the trained model and its deployed form, with minimal decision changes between the INT8 graph and the final artifact.

By Rana Muhammad Ahmed, Sabahat Abbas
Hugging Face Trending Papers
Jun 4

GMBFormer: An NDVI-Guided Global Memory Bank Transformer for Urban Green-Space Extraction from Ultra-High-Resolution Imagery

Urban green-space extraction from ultra-high-resolution (UHR) imagery is commonly performed patch by patch, which limits semantic reuse among spatially separated but visually similar vegetation patterns. Directly injecting the Normalized Difference Vegetation Index (NDVI) into red-green-blue (RGB) backbones can also blur the roles of visual appearance learning and physical vegetation confidence.

arXiv Computer Vision
Sep 3

DWFF-Net: A Multi-Scale Farmland System Habitat Identification Method with Adaptive Dynamic Weight Feature Fusion

The paper introduces DWFF‑Net, a Dynamic Weighted Feature Fusion Network designed to improve multi‑scale segmentation for agricultural habitat recognition. It employs a frozen DINOv3 encoder, a data‑level adaptive dynamic weighting strategy, and a decoder with a dynamic weight calculation network and hybrid loss. Experiments on an agricultural habitat dataset show significant gains in mIoU and mF1 over static fusion and several baseline models, especially for tiny features such as scattered trees.

By Kesong Zheng, Zhi Song, Peizhou Li, Shuyi Yao, Tong Li, Yonglin Shen, Zhenxing Bian
arXiv Computer Vision
Sep 3

PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation

PlantC2USeg is a deep transfer‑learning framework that uses cross‑scale consistency learning and an information‑restricted decoder to improve plant point cloud segmentation. It achieves state‑of‑the‑art performance on Soybean3D and ShapeNet Part, and demonstrates strong few‑shot generalization across species and sensing conditions. The method reduces the need for large annotated datasets and lowers adaptation overhead for new plant species.

By Yu Tian, Xintong Jiang, Jan Franklin Adamowski, Shiv O. Prasher, Shangpeng Sun
arXiv AI
Jul 7

SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests

arXiv:2510. 09458v2 Announce Type: replace-cross Abstract: Interest in forestry automation is growing alongside rapid advances in deep learning.

By David-Alexandre Duclos, William Guimont-Martin, Gabriel Jeanson, Arthur Larochelle-Tremblay, Martine Lapointe, Th\'eo Defosse, Fr\'ed\'eric Moore, Philippe Nolet, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
arXiv Computer Vision
4d ago

DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests

arXiv:2606.20223v2 Announce Type: replace Abstract: Camera-trap monitoring in African tropical forests increasingly extends beyond closed-canopy interiors to riverbanks, clearings, and park edges. Am...

By Hugo Magaldi, Theau d'Audiffret, Etienne Francois Akomo-Okoue, Bala Amarasekaran, Naomi Anderson, Claire Auger, Noemie Cappelle, Daniel Cornelis, Raphael Cornette, Tobias Deschner, Gabriel Dubus, Davy Fonteyn, Rosa M. Garriga, Jennifer Hatlauf, Innocent Kasekendi, Raymond Katumba, Aram Kazandjian, Alfred Ngomanda, Stephan Ntie, Simone Pika, Xavier Rufray, Harold Rugonge, John Justice Tibesigwa, Peter van Lunteren, Hadrien Vanthomme, Joeri A. Zwerts, Sabrina Krief
arXiv Machine Learning
Sep 22

Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea

The study benchmarks Vision Transformers (ViTs) against convolutional neural networks (CNNs) for fine‑grained orchid genus identification in New Guinea’s species‑rich, data‑poor flora. Using a two‑stage system that first predicts genus and then retrieves similar species images, the authors fine‑tuned four pretrained backbones on 16,701 photographs from 120 genera and 1,350 species. The self‑supervised ViT DINOv2 achieved the highest genus accuracy (macro top‑1 66.9 %) and outperformed both CNNs and a domain‑matched pretrained ViT, demonstrating strong species retrieval and open‑set detection capabilities.

By Reza Saputra, Diah Harnoni Apriyanti, Andr\'e Schuiteman, Kurt Metzger, Ashley Field, Katharina Nargar, William Edwards
arXiv Machine Learning
Sep 16

From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation

The study evaluates the use of frozen geospatial foundation embeddings (AlphaEarth) for mapping cultivated versus non‑cultivated land in Maine. Using 192 spatially separated patches and USDA Cropland Data Layer labels, a lightweight classifier achieved 93.7% overall accuracy without fine‑tuning, and a nearest‑class‑centroid rule reached 90.2%. A balanced sample of 60,000 labeled pixels was nearly as effective as the full 8.6 million‑pixel pool, and classifiers trained in one year remained accurate across 2018‑2023. In a blind human validation of 385 points, the AlphaEarth‑plus‑random‑forest map matched 95.3% of the consensus, outperforming the CDL reference (91.7%).

By Mohammad Ammar Mughees, Giovanni Montefoschi, Zhongxin Chen, Maria Antonia Brovelli