arXiv:2605. 16716v4 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
By Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat
Sanyu Studio is a multi‑agent dialogue system that treats 321 Sanyu oil paintings as agents equipped with fact, interpretation, organization, and memory‑filtering mechanisms. The paper reports on a seven‑day workshop with eight art‑university participants, showing that user prompts, evidence organization, and cognitive tendencies produced divergent yet coherent digital narratives of Sanyu. The study suggests that, when historical evidence is limited, AI can amplify human agency and provide public audiences with an interactive entry point into art‑historical interpretation.
By Zhaoxi Wei, Hongye Yang, Shuyuan Tian
arXiv:2605. 16716v5 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
By Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat
arXiv:2607. 09403v1 Announce Type: new Abstract: Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation.
By Jingbo Chen, He Wang, Wei Yuan, Yuqiao Lai, Zhenyan Lu
Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their...
arXiv:2608. 05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images.
By Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, Yixia Zheng, Peiyi Jiang, Wenjie Ning, Wei Zhang, Wei Chen
arXiv:2608.30498v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduct...
By Qi Li, Zhaojie Kang, Yingjie He, Zheng Lin, Hao Zhang, Guangxin Wu, Yan Gong, Rong Fu, Jianyuan Ni
MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.
By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse linguistic and cultural backgrounds. Therefore, adapting memes acr...
TransMeme introduces a multi‑agent framework for cross‑cultural meme transcreation, addressing the unique challenges of preserving intent, adapting cultural meaning, and maintaining multimodal consistency. The system coordinates specialized agents for cultural adaptation, text rewriting, revision, and visual adjustment, and is evaluated on Chinese‑English meme pairs. Human and LLM‑based evaluations show that TransMeme outperforms baselines, achieving a 33.1% average improvement in human scores and a 60% Top‑1 ranking rate in LLM judgments.
By Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He
arXiv:2606. 09846v1 Announce Type: cross Abstract: Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork.
By Vignesh Nagarajan
arXiv:2608. 11907v2 Announce Type: replace-cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge.
By Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang