arXiv AI

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

arXiv AI
Sep 4

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

CulturalMenuBench is a new benchmark comprising 4,870 culinary items in 10 languages across 18 regions, designed to test multimodal language models on tasks that combine dish recognition, step-by-step cooking images, ingredients, procedural text, and regional labels. The benchmark reveals a large knowledge‑application gap: models that score over 94% on standard multiple‑choice questions fall to at most 56% when attributing dishes to Chinese regional cuisines, indicating that cultural knowledge is present but not activated by visual input. Diagnostic analyses show that accuracy is driven by visual distinctiveness rather than cultural structure, and that removing sequential cooking images selectively harms process‑grounded tasks, confirming the need for procedural evidence.

By Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua Luo, Qinggang Zhang, Jinsong Su
arXiv AI
Sep 4

Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes

The paper investigates how well large language models (LLMs) can assess whether recipes are suitable for people with diabetes. It introduces a benchmark of 7,607 recipes, split evenly between suitable and unsuitable, and tests three prompting strategies that vary in how much diabetes dietary guidance they provide. Results show that LLMs tend to be cautious in labeling recipes as suitable, and those that can reason with dietary guidelines—particularly Mistral‑7B and Llama‑70B—perform best.

By Revathy Venkataramanan, Aditya Luthra, Venkatesan Nadimuthu, Amit Sheth
Hugging Face Trending Papers
Sep 3

Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes

The paper examines how large language models (LLMs) can assess whether recipes are suitable for people with diabetes. It introduces a benchmark of 7,607 recipes, split evenly between suitable and unsuitable, and tests three prompting strategies that embed varying levels of diabetes dietary guidelines. Results show that LLMs tend to be cautious in labeling recipes as suitable, and those that can reason with the guidelines—particularly Mistral‑7B and Llama‑70B—perform best.

arXiv AI
Aug 17

From Field Data to Global Food Systems Intelligence: A Semantic Graph Framework for Sustainable Wheat Production

arXiv:2502. 19507v2 Announce Type: replace Abstract: In response to the growing need for structured, interoperable agricultural data, this paper presents the Sustainable Wheat Production Datahub, a modular, graph-based framework that brings diverse wheat production datasets together into a single, queryable store.

By Nirmal Gelal, Aastha Gautam, Soheil Abadifard, Nico Giordano, Moumita Sen Sarma, Sanaz Saki Norouzi, Claudio Dias da Silva Jr, Jean Ribert Francois, Kathleen M. Jagodnik, Katherine Nelson, Terry Griffin, Xiaomao Lin, Stacy Hutchinson, Stephen M. Welch, Kelsey Andersen Onofre, Romulo Lollato, Pascal Hitzler, Hande K\"u\c{c}\"uk McGinty
arXiv AI
Jul 14

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

arXiv:2607. 10212v1 Announce Type: new Abstract: Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance.

By Nipun Misra, Vikranth Udandarao, Aanchal Gupta, Yogender Kumar, Manuj Mukherjee, Raghava Mutharaju