arXiv AI

VisTCP: A Visualization Framework to Construct Knowledge-Graph-Based Representation for Traditional Chinese Painting

arXiv:2607. 05841v1 Announce Type: cross Abstract: Structured representation can characterize semantic objects and relationships in images.

arXiv AI
Aug 7

ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

arXiv:2608. 05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images.

By Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang
arXiv Computation and Language
Sep 2

ExpArt-KG: Artwork Image Description Generation through Iterative Exploration of Knowledge Graphs

The paper introduces ExpArt-KG, a knowledge graph tailored to the artwork domain, and a retrieval‑augmented generation framework that alternates between generating answers and retrieving relevant facts from the graph. By using a correctness judgment to guide the search, the method efficiently gathers the necessary factual information, improving the detail of image explanations while reducing external knowledge retrieval costs. Experimental results demonstrate that the approach maintains generation quality comparable to fixed‑iteration methods.

By Yuta Kato, Shintaro Ozaki, Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito, Katsuhiko Hayashi, Taro Watanabe
arXiv Computer Vision
Sep 1

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

arXiv:2607.02290v2 Announce Type: replace Abstract: Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a kno...

By Zhaokai Wang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Yiguo He, Mohan Zhang, Leyao Gu, Yan Li, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Xuanhe Zhou, Zhihang Zhong, Xue Yang
arXiv AI
Aug 6

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

arXiv:2608. 05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images.

By Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, Yixia Zheng, Peiyi Jiang, Wenjie Ning, Wei Zhang, Wei Chen
arXiv AI
Sep 3

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

DocHop is a new benchmark that tests multimodal large language models on integrated chart‑context reasoning within document‑style images. The benchmark presents narrative text that imposes multi‑step compositional constraints, while charts supply the data needed to answer questions grounded in semantic reference labels. It contains 2,074 examples across six task categories, generated via a stochastic logic‑first pipeline that controls reasoning depth and visual density, and shows a large performance gap between humans (over 90% accuracy) and the best models (62.83%).

By Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee