arXiv AI By Saransh Kumar Gupta, Armaan Shah, Lipika Dey, Partha Pratim Das, Ramesh Jain

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 4

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

CulturalMenuBench is a new benchmark comprising 4,870 culinary items in 10 languages across 18 regions, designed to test multimodal language models on tasks that combine dish recognition, step-by-step cooking images, ingredients, procedural text, and regional labels. The benchmark reveals a large knowledge‑application gap: models that score over 94% on standard multiple‑choice questions fall to at most 56% when attributing dishes to Chinese regional cuisines, indicating that cultural knowledge is present but not activated by visual input. Diagnostic analyses show that accuracy is driven by visual distinctiveness rather than cultural structure, and that removing sequential cooking images selectively harms process‑grounded tasks, confirming the need for procedural evidence.

By Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua Luo, Qinggang Zhang, Jinsong Su
arXiv AI
Sep 4

Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes

The paper investigates how well large language models (LLMs) can assess whether recipes are suitable for people with diabetes. It introduces a benchmark of 7,607 recipes, split evenly between suitable and unsuitable, and tests three prompting strategies that vary in how much diabetes dietary guidance they provide. Results show that LLMs tend to be cautious in labeling recipes as suitable, and those that can reason with dietary guidelines—particularly Mistral‑7B and Llama‑70B—perform best.

By Revathy Venkataramanan, Aditya Luthra, Venkatesan Nadimuthu, Amit Sheth