arXiv Machine Learning By Andrea Morales-Garz\'on, Salvador L\'opez-Joya, Miguel L\'opez-P\'erez, Maria J. Martin-Bautista

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

Read the original on arXiv Machine Learning →

The paper studies how prompt quality affects vision‑language models in the agrifood domain. It evaluates Zero‑Shot Prompt Ensembling (ZPE) on CLIP and SigLIP across four datasets, showing that ZPE offers limited gains on in‑distribution data but significantly improves accuracy and calibration when the domain shifts. The authors also introduce Prompt‑based Inconsistency Detection (PID), which uses prompt disagreement to detect failures under severe domain shift, outperforming standard confidence measures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 9

TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases

To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation.