arXiv AI By Zhendong Li, Lei Sun, Ruibo Ming, He Zhang, Danda Pani Paudel, Luc Van Gool, Jinjin Gu

Knowledge-Centric Agents for Workflow Generation

Read the original on arXiv AI →

arXiv:2607. 15845v1 Announce Type: new Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 9

Explain Before You Answer: A Survey on Compositional Visual Reasoning

arXiv:2508. 17298v3 Announce Type: replace-cross Abstract: Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground intermediate concepts, and perform multi-step logical inference.

By Fucai Ke, Joy Hsu, Zhixi Cai, Zixian Ma, Xin Zheng, Xindi Wu, Sukai Huang, Weiqing Wang, Pari Delir Haghighi, Gholamreza Haffari, Ranjay Krishna, Jiajun Wu, Hamid Rezatofighi
arXiv Computation and Language
6d ago

Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward

The paper introduces UniSandbox, a decoupled evaluation framework with controlled synthetic datasets, to study whether understanding informs generation in Unified Multimodal Models. Results show a notable understanding‑generation gap, especially in reasoning generation and knowledge transfer. Explicit Chain‑of‑Thought (CoT) in the understanding module bridges this gap, and self‑training can internalize CoT for implicit reasoning during generation; query‑based architectures also exhibit latent CoT‑like properties that aid knowledge transfer.

By Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin, Bin Lin, Zongjian Li, Bin Zhu, Weihao Yu, Li Yuan
Hugging Face Trending Papers
Jul 21

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation.

arXiv Machine Learning
Sep 11

A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

The paper introduces a multi-stage rule‑chaining framework for the Abstraction and Reasoning Corpus (ARC), aiming to model cognitive generalization by inferring abstract rules from few examples. It combines three solvers—a deterministic rule discovery module, a pattern‑composition engine, and a structural abstraction layer—executed sequentially in a fallback hierarchy that reuses earlier reasoning traces to improve interpretability and generalization. The system achieved over 95% accuracy on ARC tasks, demonstrating strong performance across deterministic, compositional, and abstract categories.

By Deblina Kar