arXiv Computer Vision By David Ahmedt-Aristizabal, Mohammad Ali Armin, Lars Petersson

InfoTaxa: Information-Calibrated Label-Free Clustering for Fine-Grained Visual Taxonomy

Read the original on arXiv Computer Vision →

InfoTaxa presents an information‑calibrated, label‑free clustering approach for fine‑grained visual taxonomy, using frozen pretrained visual embeddings and DNA as an audit signal. On the BIOSCAN‑5M dataset, the method achieves 0.79 AMI at family and 0.67 at genus, outperforming prior image baselines and matching oracle‑K and graph‑based methods. The study shows that while clustering efficiency recovers most image‑available information at higher taxonomic ranks, species‑level performance remains limited by both clustering and representation, with DNA adding significant predictive value.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 12

Can Edge-Deployable Vision-Language Models Identify Species?

arXiv:2609. 11916v1 Announce Type: new Abstract: Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification.

By William Zhou, Mayukha Siripuram, Xiao Yan, Ziqi Liu, Yi Ding
arXiv Machine Learning
Sep 22

Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea

The study benchmarks Vision Transformers (ViTs) against convolutional neural networks (CNNs) for fine‑grained orchid genus identification in New Guinea’s species‑rich, data‑poor flora. Using a two‑stage system that first predicts genus and then retrieves similar species images, the authors fine‑tuned four pretrained backbones on 16,701 photographs from 120 genera and 1,350 species. The self‑supervised ViT DINOv2 achieved the highest genus accuracy (macro top‑1 66.9 %) and outperformed both CNNs and a domain‑matched pretrained ViT, demonstrating strong species retrieval and open‑set detection capabilities.

By Reza Saputra, Diah Harnoni Apriyanti, Andr\'e Schuiteman, Kurt Metzger, Ashley Field, Katharina Nargar, William Edwards
arXiv Computer Vision
Sep 1

Automated pipeline for herbarium label digitization

HERBIOME is a modular, end‑to‑end pipeline that automates the digitization of herbarium labels. It combines YOLOv8 for component detection, CRAFT Hezar for word‑level text localization, a fine‑tuned TrOCR model for mixed handwritten and printed text recognition, and GPT‑4o Mini for structuring metadata into standardized fields. Evaluation on 450 French specimens shows high surface similarity (MWS ≈ 0.616) and moderate semantic accuracy (SMA ≈ 0.442), with taxonomic fields identified as the main challenge.

By Hiba Abbad, Hanane Ariouat, Eva Perez Pimpare, Nicolas Turenne, Eric Chenin, Abderrazak Sebaa, Edi Prifti, Jean-Daniel Zucker, Youcef Sklab
arXiv Computer Vision
4d ago

DeepForestVisionV2: Ecology-Driven Taxonomy Expansion for Camera-Trap Monitoring in African Tropical Forests

arXiv:2606.20223v2 Announce Type: replace Abstract: Camera-trap monitoring in African tropical forests increasingly extends beyond closed-canopy interiors to riverbanks, clearings, and park edges. Am...

By Hugo Magaldi, Theau d'Audiffret, Etienne Francois Akomo-Okoue, Bala Amarasekaran, Naomi Anderson, Claire Auger, Noemie Cappelle, Daniel Cornelis, Raphael Cornette, Tobias Deschner, Gabriel Dubus, Davy Fonteyn, Rosa M. Garriga, Jennifer Hatlauf, Innocent Kasekendi, Raymond Katumba, Aram Kazandjian, Alfred Ngomanda, Stephan Ntie, Simone Pika, Xavier Rufray, Harold Rugonge, John Justice Tibesigwa, Peter van Lunteren, Hadrien Vanthomme, Joeri A. Zwerts, Sabrina Krief