arXiv Machine Learning

Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea

The study benchmarks Vision Transformers (ViTs) against convolutional neural networks (CNNs) for fine‑grained orchid genus identification in New Guinea’s species‑rich, data‑poor flora. Using a two‑stage system that first predicts genus and then retrieves similar species images, the authors fine‑tuned four pretrained backbones on 16,701 photographs from 120 genera and 1,350 species. The self‑supervised ViT DINOv2 achieved the highest genus accuracy (macro top‑1 66.9 %) and outperformed both CNNs and a domain‑matched pretrained ViT, demonstrating strong species retrieval and open‑set detection capabilities.

arXiv Computer Vision
Sep 3

PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation

PlantC2USeg is a deep transfer‑learning framework that uses cross‑scale consistency learning and an information‑restricted decoder to improve plant point cloud segmentation. It achieves state‑of‑the‑art performance on Soybean3D and ShapeNet Part, and demonstrates strong few‑shot generalization across species and sensing conditions. The method reduces the need for large annotated datasets and lowers adaptation overhead for new plant species.

By Yu Tian, Xintong Jiang, Jan Franklin Adamowski, Shiv O. Prasher, Shangpeng Sun
arXiv AI
Jul 7

SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests

arXiv:2510. 09458v2 Announce Type: replace-cross Abstract: Interest in forestry automation is growing alongside rapid advances in deep learning.

By David-Alexandre Duclos, William Guimont-Martin, Gabriel Jeanson, Arthur Larochelle-Tremblay, Martine Lapointe, Th\'eo Defosse, Fr\'ed\'eric Moore, Philippe Nolet, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
arXiv AI
Aug 24

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

AT‑ViT is a dual‑branch Vision Transformer that processes both raw herbarium scans and their segmentation masks through a multi‑scale, multi‑view cross‑attention fusion. It uses a mask‑guided patch weighting scheme to emphasize plant regions and suppress background artifacts, thereby encouraging plant‑centric representations. In trait classification tasks such as leaf base shape and thorns, AT‑ViT consistently outperforms baselines, improves spatial attention grounding (IoU_p +15.66 to +18.03 pp, IoU_b –27.92 to –31.02 pp), and shows greater robustness to synthetic background perturbations, surpassing ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points. whyItMatters":"The model addresses shortcut learning caused by background cues in herbarium images, leading to more accurate and interpretable plant trait recognition."

By Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti
arXiv AI
Jun 2

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

arXiv:2606. 02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural management.

By Adri\'an C\'anovas-Rodriguez, Miguel A. Gonz\'alez-Ill\'an, Maria Fernanda Garc\'ia-Cruz, Pedro Nortes Tortosa, Jos\'e Salvador Rubio-Asensio, Miguel A. Zamora Izquierdo, Juan Antonio Mart\'inez Navarro, Antonio F. Skarmeta
arXiv Computer Vision
Sep 25

LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping

LeafTrackNet is a deep learning framework that combines a YOLOv10-based leaf detector with a MobileNetV3-based embedding network to track individual leaves over time. The authors introduce CanolaTrack, a large benchmark dataset of 5,704 RGB images with 31,840 annotated leaf instances from 184 canola plants. When evaluated without prior fine‑tuning, LeafTrackNet outperforms existing methods on CanolaTrack, KOMATSUNA, and MSU‑PID datasets, achieving HOTA scores of 88.03, 87.33, and 74.20 respectively.

By Shanghua Liu, Majharulislam Babor, Christoph Verduyn, Breght Vandenberghe, Bruno Betoni Parodi, Cornelia Weltzien, Marina M. -C. H\"ohne
arXiv Machine Learning
Aug 27

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

CropCop is a closed‑set plant‑health recognition system covering 120 operational classes, built from a rigorously audited dataset of 109,107 images after removing 3,233 duplicate relationships. The model, based on a fine‑tuned DINOv3 ConvNeXt‑Tiny, achieves 98.51% accuracy and 96.87% macro‑F1 on a locked internal test, while a quantised MobileNetV4 variant reaches 98.46% accuracy and 96.23% macro‑F1 in a 22.60 MiB runtime artifact. Validation‑only post‑training quantisation and a compact ExecuTorch/XNNPACK PTE ensure high fidelity between the trained model and its deployed form, with minimal decision changes between the INT8 graph and the final artifact.

By Rana Muhammad Ahmed, Sabahat Abbas
arXiv AI
Aug 3

Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping

arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.

By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
arXiv Computer Vision
Sep 17

CoAtNet-DeepMoE: A Convolution-Attention Hybrid with DeepSeek Mixture-of-Experts for Parameter-Efficient Tomato Disease Classification

CoAtNet-DeepMoE is a lightweight Convolution‑Attention hybrid architecture that incorporates a DeepSeek Mixture‑of‑Experts to reduce parameters while maintaining high accuracy for tomato disease classification. The model achieves state‑of‑the‑art performance on Kaggle and PlantVillage datasets, reporting 99.80% accuracy on Kaggle and 99.83% accuracy on PlantVillage, all with only 2.47 million parameters. The source code will be released on GitHub.

By Md Nadim Mahamood, Md Arif Shahriar, Md Shafi Ud Doula, Kamrul Hasan