Manga visual question answering requires models to answer questions over panel-based visual narratives, where relevant evidence is distributed across ordered panels, embedded text, recurring character...
ManGo is an unsupervised framework for manga visual question answering that actively selects panels, extracts concise clues, and decides when to stop, creating a compact evidence sketch before answering. It introduces Active Narrative Sketching (ANS) and optimizes its behavior using group-relative policy training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. Experiments on standard manga understanding benchmarks demonstrate that ManGo achieves state‑of‑the‑art performance across different settings.
By Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao
SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.
By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
arXiv:2602.17690v3 Announce Type: replace-cross
Abstract: Graphic design generation demands a delicate balance between high visual fidelity and fine-grained structural editability. However, existing...
By Ziyuan Liu, Shizhao Sun, Danqing Huang, Yingdong Shi, Meisheng Zhang, Ji Li, Jingsong Yu, Jiang Bian
Editable Visual Design introduces a new design paradigm that combines a Coding Agent with a Vision‑Language Model (VLM) and an image generation model. The VLM acts as the creative brain, understanding requirements, planning tasks, and judging aesthetics, while the image generator produces isolated visual assets on demand. The agent follows an "imagine first, then act" workflow, generating assets, writing native HTML/CSS, and refining the design through visual feedback, ultimately producing editable, layer‑wise artifacts with real text that can be adjusted via a graphical interface.
By Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen, HaoDong Li, Zhutao Lv, Jiaxin Lin, Jinhua Yu, Jun He, Zilong Huang, Rui Chen, Weijia Li
arXiv:2608. 03298v1 Announce Type: new Abstract: Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usable artifacts.
By Shengjun Fang, Chenyang Wu, Zongzhang Zhang
arXiv:2608. 16289v3 Announce Type: replace Abstract: Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs.
By Xiaoan Liu, Lichen Ma, Zipeng Guo, Yu He, Xiaoyan Su, Shaojie Guo, Jingling Fu, Xiaolong Fu, Hao Yang, Tongxuan Liu, Yu Guo, Fei Wang, Xinyi Liu, Yongjun Zhang, Junshi Huang
arXiv:2507. 17853v2 Announce Type: replace-cross Abstract: Recent advances in text-to-image (T2I) generation have led to impressive visual results.
By Lifeng Chen, Jiner Wang, Zihao Pan, Beier Zhu, Xiaofeng Yang, Chi Zhang
arXiv:2609.36593v1 Announce Type: cross
Abstract: Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be des...
By Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li
arXiv:2607. 19038v1 Announce Type: cross Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives.
By Jialong Zuo, Haotong Zuo, Shiwei Zhang, Xiang Wang, Chen Li, Nong Sang, Changxin Gao, Xiang Bai
arXiv:2606. 26171v1 Announce Type: cross Abstract: Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequential outputs, as required in comics, storyboards, and visual narratives.
By Zihao Wang, Yijia Xu, Haoze Zheng, Xuran Ma, Haokun Gui, Harry Yang
arXiv:2608. 10337v1 Announce Type: cross Abstract: We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of narrative constraints at selected moments in a story, then use AI to generate intervening prose.
By Chao Zhang, Abe Davis