arXiv AI

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

arXiv:2608. 03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes.

arXiv AI
Jun 9

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

arXiv:2606. 08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large multimodal datasets linking diverse foods to complete nutrient profiles.

By Runze Yan, Minxiao Wang, Jiaying Lu, Darren Liu, Xiao Hu, Hanqi Luo
arXiv AI
Sep 4

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

CulturalMenuBench is a new benchmark comprising 4,870 culinary items in 10 languages across 18 regions, designed to test multimodal language models on tasks that combine dish recognition, step-by-step cooking images, ingredients, procedural text, and regional labels. The benchmark reveals a large knowledge‑application gap: models that score over 94% on standard multiple‑choice questions fall to at most 56% when attributing dishes to Chinese regional cuisines, indicating that cultural knowledge is present but not activated by visual input. Diagnostic analyses show that accuracy is driven by visual distinctiveness rather than cultural structure, and that removing sequential cooking images selectively harms process‑grounded tasks, confirming the need for procedural evidence.

By Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua Luo, Qinggang Zhang, Jinsong Su
arXiv Computer Vision
Sep 17

CoAtNet-DeepMoE: A Convolution-Attention Hybrid with DeepSeek Mixture-of-Experts for Parameter-Efficient Tomato Disease Classification

CoAtNet-DeepMoE is a lightweight Convolution‑Attention hybrid architecture that incorporates a DeepSeek Mixture‑of‑Experts to reduce parameters while maintaining high accuracy for tomato disease classification. The model achieves state‑of‑the‑art performance on Kaggle and PlantVillage datasets, reporting 99.80% accuracy on Kaggle and 99.83% accuracy on PlantVillage, all with only 2.47 million parameters. The source code will be released on GitHub.

By Md Nadim Mahamood, Md Arif Shahriar, Md Shafi Ud Doula, Kamrul Hasan
arXiv AI
Jun 9

AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning

arXiv:2603. 14342v2 Announce Type: replace-cross Abstract: Modern agricultural data is sourced from diverse platforms and spans multiple spatial scales, ranging from ground-level close-up photography to Unmanned Aerial Vehicle (UAV) aerial observation and satellite remote sensing imagery.

By Jiarui Zhang, Junqi Hu, Zurong Mai, Yang Liu, Yuhang Chen, Shuohong Lou, Henglian Huang, Hong Cheng, Lingyuan Zhao, Jianxi Huang, Yutong Lu, Haohuan Fu, Juepeng Zheng
Hugging Face Trending Papers
Aug 9

TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases

To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation.