arXiv AI

WhiteTesseract: Reframing the Interpretation of Cultural Heritage through XR and Conversational AI

arXiv:2605. 16972v2 Announce Type: replace-cross Abstract: Cultural heritage exhibitions often struggle to sustain attention and support reflective engagement.

arXiv Computer Vision
Sep 10

The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library

The Living Library is an end‑to‑end framework that converts fragmented digital archives into governed, conversational exhibit experiences. Developed at the Theodore Roosevelt Presidential Library, it digitizes a 300,000‑record collection, enriches it with OCR and metadata, and publishes it to a hybrid dense/semantic index. The system supports curator review via the Archivist App, powers a researcher interface, and runs Talk to TR—a museum exhibit where a digital human embodiment of Theodore Roosevelt answers visitors’ questions using Cross‑Era Analogical Grounding and dual‑path retrieval to keep responses grounded and responsive.

By Pengce Wang, Lucia Ronchi Darre, Matt Briney, Michaell Bakalars, Dan Rutkowski, Ursula Hardy, David Wolf, Laura Hoffman, Allen Kim, Shawn Wright, Juan Lavista Ferres
arXiv AI
Jun 30

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.

By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
arXiv AI
Sep 17

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.

By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
arXiv AI
Aug 6

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

arXiv:2608. 05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images.

By Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, Yixia Zheng, Peiyi Jiang, Wenjie Ning, Wei Zhang, Wei Chen
arXiv AI
Sep 10

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

NormViz introduces a new benchmark, NormViz‑Bench, comprising 3,268 contrastive image pairs from 16 countries that test AI’s ability to recognize culturally relevant visual norms. Each pair differs only in a behavior that changes its cultural interpretation, and images are labeled as conforming, violating, or irrelevant to local norms, requiring both images to be correctly classified. The benchmark shows current VLMs perform poorly, and a complementary training set, NormViz‑Train, offers a path to improve performance by teaching models to link visual perception with cultural significance.

By Akhila Yerukola, Fabrice Y Harel-Canada, Simran Khanuja, Abhinav Sukumar Rao, Ashima Suvarna, Nanyun Peng, Saadia Gabriel, Maarten Sap
arXiv AI
Jun 18

Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals

arXiv:2505. 16057v2 Announce Type: replace-cross Abstract: AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have significantly lowered the technical barriers for producing highly realistic audio, images, and videos through simple natural language prompts.

By Ayae Ide, Tory Park, Jaron Mink, Tanusree Sharma