arXiv AI

Geometry-Enhanced Portion Estimation for Multimodal LLMs

arXiv:2607. 16514v1 Announce Type: cross Abstract: Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker.

arXiv AI
Aug 5

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

arXiv:2608. 03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes.

By Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis
arXiv AI
Jun 9

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

arXiv:2606. 08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large multimodal datasets linking diverse foods to complete nutrient profiles.

By Runze Yan, Minxiao Wang, Jiaying Lu, Darren Liu, Xiao Hu, Hanqi Luo
arXiv Computer Vision
Sep 17

CALIPER: Metric-Grounded Model-Free Recognition of Visually Similar Industrial Parts

CALIPER is a model‑free RGB‑D framework that performs fine‑grained recognition of visually similar industrial parts by combining support‑based appearance matching with metric size evidence. Each class is onboarded from a single turntable RGB‑D video and a few labeled real images, enabling 3D reconstruction for appearance support and depth‑aligned size profiling. At inference, a YOLOv8n‑seg model localizes parts, a frozen DINOv2 backbone with an episodically trained embedding head matches support, and margin‑conditioned metric fusion selectively uses size evidence for ambiguous cases, achieving high accuracy on 18 parts and robust enrollment of unseen screws without retraining.

By Alankrit Gupta, Chenxi Tao, Seung-Kyum Choi
Hugging Face Trending Papers
Jul 22

Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data

Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing both visual growth dynamics and per-plant measurement labels are scarce. In this paper, we introduce a novel, annotated image time-series dataset of 691 sweet pepper plants monitored over two growing seasons, comprising 4837 images with per-plant fruit counts categorized by maturity.

arXiv Computer Vision
Sep 3

PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation

PlantC2USeg is a deep transfer‑learning framework that uses cross‑scale consistency learning and an information‑restricted decoder to improve plant point cloud segmentation. It achieves state‑of‑the‑art performance on Soybean3D and ShapeNet Part, and demonstrates strong few‑shot generalization across species and sensing conditions. The method reduces the need for large annotated datasets and lowers adaptation overhead for new plant species.

By Yu Tian, Xintong Jiang, Jan Franklin Adamowski, Shiv O. Prasher, Shangpeng Sun
arXiv AI
Sep 23

Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information

The study introduces a multimodal framework that predicts postprandial glycemic response (PPGR) by combining image-derived macronutrient estimates with clinical variables and gut microbiome data. It jointly performs macronutrient estimation from meal images and glucose prediction, using an attention-based module to model interactions between dietary and host-specific information. Evaluated on a real-world dataset, the model outperforms existing PPGR baselines that use image-derived inputs and nearly matches methods relying on manually reported macronutrients.

By Varvara Kondratyeva, Kamilia Zaripova, Nassir Navab, Azade Farshad
arXiv Machine Learning
Sep 17

MCLC-NET: Multimodal Continual Learning for Leaf Counting

The paper introduces MCLC‑NET, a multimodal continual learning framework for leaf counting that sequentially learns tasks using a memory buffer to retain key samples. It also presents MMLC, a new real‑world dataset containing RGB, depth, and thermal images across different crops and environmental conditions, organized in crop‑wise, time‑wise, and mixed orderings. Experiments show that MCLC‑NET outperforms existing methods on all three task orderings, achieving the lowest average mean squared errors.

By Ruchi Bhatt, Pratibha Kumari, Shreya Bansal, Vedant Agnihotri, Dwarikanath Mahapatra, Mukesh Saini