arXiv Machine Learning

ViFA-Council: Multi-Agent LLM Deliberation for Vietnamese Folk Art Generation

arXiv AI
Aug 20

Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction

Sanyu Studio is a multi‑agent dialogue system that treats 321 Sanyu oil paintings as agents equipped with fact, interpretation, organization, and memory‑filtering mechanisms. The paper reports on a seven‑day workshop with eight art‑university participants, showing that user prompts, evidence organization, and cognitive tendencies produced divergent yet coherent digital narratives of Sanyu. The study suggests that, when historical evidence is limited, AI can amplify human agency and provide public audiences with an interactive entry point into art‑historical interpretation.

By Zhaoxi Wei, Hongye Yang, Shuyuan Tian
arXiv AI
Aug 6

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

arXiv:2608. 05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images.

By Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, Yixia Zheng, Peiyi Jiang, Wenjie Ning, Wei Zhang, Wei Chen
arXiv AI
Sep 17

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.

By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
arXiv AI
Aug 28

TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation

TransMeme introduces a multi‑agent framework for cross‑cultural meme transcreation, addressing the unique challenges of preserving intent, adapting cultural meaning, and maintaining multimodal consistency. The system coordinates specialized agents for cultural adaptation, text rewriting, revision, and visual adjustment, and is evaluated on Chinese‑English meme pairs. Human and LLM‑based evaluations show that TransMeme outperforms baselines, achieving a 33.1% average improvement in human scores and a 60% Top‑1 ranking rate in LLM judgments.

By Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He