COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows
arXiv:2607. 01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently.
arXiv:2606. 00931v1 Announce Type: cross Abstract: Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture the diversity of real-image tasks in professional workflows.
arXiv:2607. 01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently.
VBVR-Pro is a closed‑loop testbed that enables native visual reasoning through generation, offering 300 procedurally generated tasks that scale training and allow strong transfer to external benchmarks. It supplies verifiable reward scorers based on deterministic, task‑specific rules, outperforming VLM‑as‑a‑judge approaches and providing reliable signals for reinforcement learning. The suite also facilitates controlled modality studies, revealing that video generation excels at persistent spatiotemporal tracking while interleaved generation offers a compute‑efficient alternative, and highlights the importance of vision‑native trajectories for reasoning.
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation.
arXiv:2608. 09666v1 Announce Type: new Abstract: Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive.
arXiv:2608. 14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time.
OmniEdit-Bench introduces a comprehensive benchmark for instruction-based video editing (IVE), addressing limitations of existing datasets by covering spatial, temporal, audio, and reference-based editing tasks and distinguishing explicit from implicit instructions. The evaluation framework assesses editing quality across accuracy, preservation, realism, and consistency, using human judgments and vision-language models, and incorporates an accuracy-aware penalty to ensure instruction fidelity. Experiments reveal that current IVE models perform poorly, highlighting the need for improved methods.
arXiv:2607. 08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing.
arXiv:2606. 27608v1 Announce Type: cross Abstract: We present Qwen-Image-2.
PrismGPT is a Vision‑Language Model that generates structured, region‑aware photo‑editing plans from a single image, without relying on commercial black‑box tools. It learns to diagnose aesthetic issues globally and locally while predicting precise editing parameters, using proxy‑guided learning with operation decomposition and region‑aware aesthetic ranking to bootstrap the model. A competence‑based dynamic scheduler shifts training focus from proxy tasks to the main editing task as skills improve, and all reasoning traces for fine‑tuning are self‑synthesized by the model itself. Experiments on MIT‑Adobe FiveK and a new professionally retouched benchmark, SPIRE, show PrismGPT achieves state‑of‑the‑art results using only about 6% of the training data required by previous methods.
arXiv:2608. 09111v1 Announce Type: new Abstract: AI video generation has advanced rapidly and entered widespread commercial use.
arXiv:2602. 11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as instructional and educational media.
The paper introduces VWG-Bench, a benchmark covering nine reasoning dimensions and 38 tasks to evaluate video generative models on symbolic reasoning, physical laws, and goal pursuit. It also presents Vid-PRE, a prompt-rewriting framework that offloads reasoning to a VLM, improving logical performance without changing the generator architecture. Experiments show that current models excel at visual quality but struggle with logic-heavy tasks, while Vid-PRE significantly boosts reasoning across multiple generators.