The Living Library is an end‑to‑end framework that converts fragmented digital archives into governed, conversational exhibit experiences. Developed at the Theodore Roosevelt Presidential Library, it digitizes a 300,000‑record collection, enriches it with OCR and metadata, and publishes it to a hybrid dense/semantic index. The system supports curator review via the Archivist App, powers a researcher interface, and runs Talk to TR—a museum exhibit where a digital human embodiment of Theodore Roosevelt answers visitors’ questions using Cross‑Era Analogical Grounding and dual‑path retrieval to keep responses grounded and responsive.
By Pengce Wang, Lucia Ronchi Darre, Matt Briney, Michaell Bakalars, Dan Rutkowski, Ursula Hardy, David Wolf, Laura Hoffman, Allen Kim, Shawn Wright, Juan Lavista Ferres
arXiv:2606. 09846v1 Announce Type: cross Abstract: Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork.
By Vignesh Nagarajan
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges.
arXiv:2607. 03731v1 Announce Type: cross Abstract: Creating 3D assets for virtual reality requires modeling expertise, which restricts the authorship of immersive experiences.
By Weiwei Jiang, Wanyu He, Zheyu Tan, Zheyuan Kuang, Difeng Yu, Shinobu Hasegawa, Sven Mayer, Zhanna Sarsenbayeva
arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.
By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.
By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
arXiv:2608. 05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images.
By Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, Yixia Zheng, Peiyi Jiang, Wenjie Ning, Wei Zhang, Wei Chen
arXiv:2609.15696v1 Announce Type: cross
Abstract: Generative AI (GenAI) tools are increasingly woven into how blind and low-vision (BLV) people communicate, not only with digital information, but wit...
By Protik Dey, Mohd Saifuzzaman, Taslima Akter
NormViz introduces a new benchmark, NormViz‑Bench, comprising 3,268 contrastive image pairs from 16 countries that test AI’s ability to recognize culturally relevant visual norms. Each pair differs only in a behavior that changes its cultural interpretation, and images are labeled as conforming, violating, or irrelevant to local norms, requiring both images to be correctly classified. The benchmark shows current VLMs perform poorly, and a complementary training set, NormViz‑Train, offers a path to improve performance by teaching models to link visual perception with cultural significance.
By Akhila Yerukola, Fabrice Y Harel-Canada, Simran Khanuja, Abhinav Sukumar Rao, Ashima Suvarna, Nanyun Peng, Saadia Gabriel, Maarten Sap
arXiv:2606. 19727v1 Announce Type: cross Abstract: Language models have become essential tools in shaping modern workflows.
By Punit Kumar Singh, Niladri Ghosh, Advait Joshi{\i}nst, Shailee Choudhary, Michael F\"arber, Haiqin Yang
arXiv:2505. 16057v2 Announce Type: replace-cross Abstract: AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have significantly lowered the technical barriers for producing highly realistic audio, images, and videos through simple natural language prompts.
By Ayae Ide, Tory Park, Jaron Mink, Tanusree Sharma
arXiv:2606. 11835v1 Announce Type: cross Abstract: Collecting participants' lived experiences is central to design research.
By Zhiqing Wang, Steven Dow