arXiv Computation and Language
Sep 1

ManGo: Manga Active Narrative Grounding Optimization

ManGo is an unsupervised framework for manga visual question answering that actively selects panels, extracts concise clues, and decides when to stop, creating a compact evidence sketch before answering. It introduces Active Narrative Sketching (ANS) and optimizes its behavior using group-relative policy training with two rewards: answer preference from listwise self-ranking and path consistency from stable ordered panel trajectories. Experiments on standard manga understanding benchmarks demonstrate that ManGo achieves state‑of‑the‑art performance across different settings.

By Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao
arXiv AI
Sep 15

Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation

The paper introduces SCoRE, an agentic framework for Visual Retrieval-Augmented Generation that explicitly selects and consolidates visual evidence before generating answers. It addresses two key challenges: sparse, scattered evidence and noisy exploration trajectories that obscure reasoning. By maintaining a textual ledger of relevant observations and reloading original images for a logical evidence sequence, SCoRE decouples reasoning from exploration and enforces strict visual grounding, with training that rewards evidence coverage, compactness, and answer correctness.

By Yucheng Shen, Lingyong Yan, Jiulong Wu, Shuaiqiang Wang, Jianmin WU, Dawei Yin, Min Cao
arXiv AI
Jul 22

PlotTwist: A Creative Plot Generation Framework with Small Language Models

arXiv:2603. 16410v2 Announce Type: replace-cross Abstract: Creative plot generation presents a fundamental challenge for language models: transforming a concise premise into a coherent narrative that sustains global coherence, character development, pacing, tone consistency, and emotional progression.

By Abhinav Thorat, Ravi Kolla, Jyotin Goel, Madhav Kataria, Niranjan Pedanekar
arXiv AI
Sep 15

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

NoteVQA is a new benchmark that collects 252 real‑life visual questions from the Chinese image‑sharing platform Xiaohongshu, covering 12 topics and 7 user intents. Each question is paired with a concise expert reference and a human‑audited interleaved answer that blends text and visual evidence. The study evaluates VLMs on short‑answer correctness and interleaved answer quality using a new AgenticInterleave framework and a 12‑dimensional IVR‑12 rubric, finding that even state‑of‑the‑art models achieve only about 53% accuracy and lag behind human references in content quality.

By Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen, Yahui Liu, Yao Hu, Chuan Mu