The study introduces a multimodal framework that predicts postprandial glycemic response (PPGR) by combining image-derived macronutrient estimates with clinical variables and gut microbiome data. It jointly performs macronutrient estimation from meal images and glucose prediction, using an attention-based module to model interactions between dietary and host-specific information. Evaluated on a real-world dataset, the model outperforms existing PPGR baselines that use image-derived inputs and nearly matches methods relying on manually reported macronutrients.
By Varvara Kondratyeva, Kamilia Zaripova, Nassir Navab, Azade Farshad
arXiv:2607. 08423v1 Announce Type: new Abstract: The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management.
By Qian Jiang, Zhecheng Shi, Jingpu Yang, Zirui Song, Miao Fang
arXiv:2608. 03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes.
By Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis
arXiv:2609.22171v1 Announce Type: cross
Abstract: Precision healthcare, particularly for conditions like hypertension and cardiovascular disease, necessitates monitoring of dietary sodium intake. How...
By Mingyu Huang, Weiqing Min, Yuehui Fang, Yuna He, Shuqiang Jiang
arXiv:2607. 07673v1 Announce Type: cross Abstract: Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams.
By Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
arXiv:2606. 10120v1 Announce Type: cross Abstract: Postprandial hyperglycemia is a key risk factor for metabolic disorders; however, existing dietary guidance is often static, impractical, and insufficiently personalized, providing recommendations that are difficult to follow or not impactful.
By Asiful Arefeen, Carol Johnston, Hassan Ghasemzadeh
arXiv:2609.22099v1 Announce Type: new
Abstract: Cooking is a complex process that transforms raw ingredients into delicious and nutritious dishes, yet the recipes that encode this process remain larg...
By Mansi Goel, Sumit Bhagat, Saloni Srivastava, Malav Patel, Shlok Vinodkumar Mehroliya, Ganesh Bagler
arXiv:2607. 23273v1 Announce Type: cross Abstract: Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than automated reasoning.
By James Izzard, Hassan Eshkiki, Fabio Caraffini
arXiv:2607. 16514v1 Announce Type: cross Abstract: Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker.
By Lin Liao, Peng Li
AgriScope is a unified pixel‑grounded multimodal framework designed for agricultural image understanding. It supports image‑level, region‑level, and pixel‑level tasks such as grounded caption generation, referring expression segmentation, and multi‑turn multimodal interaction. The authors also introduce AgriGround, a large‑scale dataset with over 500K images and 11M instruction‑following samples, created via an automatic annotation pipeline that combines caption generation, phrase‑level grounding, segmentation mask creation, and instruction synthesis.
By Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain, Sajid Javed
The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.
By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach
arXiv:2605. 18419v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology.
By Franciskus Xaverius Erick, Johanna Paula M\"uller, Bernhard Kainz