ReGraph: Learning to Generate Recipe Graphs from Food Images
arXiv:2608. 06917v1 Announce Type: new Abstract: Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.
The paper introduces ViralRecipesTrans, a dataset of execution flow graphs from culinary videos linked to specific creators, and proposes a graph learning framework to discover procedural personas. It shows that discrete topological metrics better capture a creator’s workflow than lexical classifiers, and presents a two‑stage generative model that predicts a creator’s exact execution graph for new dishes. The study finds that few‑shot LLMs excel at semantic assignment but lack macro‑planning, while the structured model offers superior topological control, and an ensemble approach combines both strengths for personalized workflow generation.
arXiv:2608. 06917v1 Announce Type: new Abstract: Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.
While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.
arXiv:2508. 09860v2 Announce Type: replace Abstract: Human-aligned AI is a critical component of co-creativity, as it enables models to accurately interpret human intent and generate controllable outputs that align with design goals in collaborative content creation.
arXiv:2607. 15845v1 Announce Type: new Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions.
arXiv:2607. 00527v1 Announce Type: new Abstract: Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime.
arXiv:2607. 15845v2 Announce Type: replace Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions.
The paper introduces Forking Garden, a branching game generation system that threads narrative archetype as a persistent semantic signal throughout the generation pipeline. Narrative progression is modeled with soft Rise/Fall states, guiding the creation of plot nodes, structural constraints, encounter composition, objectives, rewards, and difficulty adaptation. Experiments across ten storylines show that the system produces distinct archetypal trajectories, improves entity diversity, and that Rise/Fall distinctions remain meaningful during play, aiding both narrative understanding and creator interpretation.
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks.
arXiv:2603. 11863v2 Announce Type: replace Abstract: The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve.
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
arXiv:2607. 09403v1 Announce Type: new Abstract: Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation.
The paper shows that large language models (LLMs) naturally organize their hidden state manifolds into small‑world networks, enabling efficient multi‑hop reasoning. By converting similarity matrices into unweighted graphs, the authors trace connectivity between distant semantic anchors and find a sharp topological phase transition: deep reasoning layers compress conceptual distances into paths bounded by six semantic hops, while early syntactic layers remain fragmented. The framework is applied to zero‑shot hallucination detection in Retrieval‑Augmented Generation, revealing that factual generations preserve a ~3‑hop structure, whereas hallucinations collapse the topology.