MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Manga visual question answering requires models to answer questions over panel-based visual narratives, where relevant evidence is distributed across ordered panels, embedded text, recurring character...
ManGo is an unsupervised framework for manga visual question answering that actively selects panels, extracts concise clues, and decides when to stop, creating a compact evidence sketch before answering. It introduces Active Narrative Sketching (ANS) and optimizes its behavior using group-relative policy training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. Experiments on standard manga understanding benchmarks demonstrate that ManGo achieves state‑of‑the‑art performance across different settings.
SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.
arXiv:2602.17690v3 Announce Type: replace-cross Abstract: Graphic design generation demands a delicate balance between high visual fidelity and fine-grained structural editability. However, existing...
Editable Visual Design introduces a new design paradigm that combines a Coding Agent with a Vision‑Language Model (VLM) and an image generation model. The VLM acts as the creative brain, understanding requirements, planning tasks, and judging aesthetics, while the image generator produces isolated visual assets on demand. The agent follows an "imagine first, then act" workflow, generating assets, writing native HTML/CSS, and refining the design through visual feedback, ultimately producing editable, layer‑wise artifacts with real text that can be adjusted via a graphical interface.
arXiv:2608. 03298v1 Announce Type: new Abstract: Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usable artifacts.