arXiv AI

ReGraph: Learning to Generate Recipe Graphs from Food Images

arXiv:2608. 06917v1 Announce Type: new Abstract: Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.

arXiv AI
Aug 26

Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs

The paper introduces ViralRecipesTrans, a dataset of execution flow graphs from culinary videos linked to specific creators, and proposes a graph learning framework to discover procedural personas. It shows that discrete topological metrics better capture a creator’s workflow than lexical classifiers, and presents a two‑stage generative model that predicts a creator’s exact execution graph for new dishes. The study finds that few‑shot LLMs excel at semantic assignment but lack macro‑planning, while the structured model offers superior topological control, and an ensemble approach combines both strengths for personalized workflow generation.

By Lei Jiang
arXiv AI
Aug 17

RecipeNet: A Hierarchical Transformer for Recipe Data

arXiv:2608. 14505v1 Announce Type: cross Abstract: Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, where procedures are represented as ordered sequences of steps containing heterogeneous structured fields.

By Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi, Abhinav Kumar, Baoxin Li
arXiv Computer Vision
Aug 31

Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning

The paper introduces Visual Metaphor Transfer (VMT), a task that requires models to extract the abstract ‘creative essence’ from a reference image and apply it to a new target subject. It proposes a multi‑agent framework based on Conceptual Blending Theory, using a Schema Grammar to separate relational invariants from visual entities. The system includes perception, transfer, generation, and diagnostic agents, and experimental results show it outperforms state‑of‑the‑art baselines in metaphor consistency, analogy appropriateness, and visual creativity.

By Yu Xu, Yuxin Zhang, Lin Gao, Oliver Deussen, Tong-Yee Lee, Fan Tang
arXiv AI
Jun 24

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

arXiv:2606. 24849v1 Announce Type: cross Abstract: Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object counts, spatial relations, attribute bindings, and coarse layouts must be preserved.

By Zixuan Li, Haokun Lin, Yicheng Xiao, Zhiwei Li, Xinyang Song, Zelong Zheng, Yong He, Heng Yao, Ke Ding, Chao Yu, Chuan Yuan, Qi Li, Zhenan Sun