Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability
arXiv:2607.
Diffusion TV is an interactive AI art installation that transforms a CRT TV into a tangible interface for diffusion models. By physically adjusting the TV’s antenna, users control the clarity of AI-generated images and sounds, symbolically mirroring the denoising process that underlies diffusion-based generation. The installation offers three channels—Past, Present, and Future—featuring AI-generated animals, allowing participants to explore intermediate states of the generative process through continuous audiovisual feedback.
arXiv:2607.
Passing is an interactive audiovisual installation that transforms a single continuous monorail-window recording into an endless journey by reconstructing it as a spatiotemporal volume and resampling its spatial and temporal structure along nonlinear trajectories. A camera-based viewer‑presence detection system influences transitions among rendered video sequences, and the resulting video stream is fed into SpecMaskFoley, a real‑time video‑to‑audio synthesis model that generates a synchronized soundscape. The work distributes creative agency among the artist, the AI model, and the audience, exploring how authorship and listening can be negotiated among human intention, machine inference, and audience interpretation.
arXiv:2607. 09705v1 Announce Type: cross Abstract: Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that progressively degrade model performance.
arXiv:2605. 16972v2 Announce Type: replace-cross Abstract: Cultural heritage exhibitions often struggle to sustain attention and support reflective engagement.
Diffusion models and flow-based models have recently become the dominant paradigms in generative modeling, largely due to their ability to learn rich, multi-level visual representations through large-...
arXiv:2601. 00664v2 Announce Type: replace-cross Abstract: Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation.
arXiv:2609.13352v1 Announce Type: cross Abstract: We present three robotic art installations which explore the aesthetics of adaptive behavior. Through embodied machine leaning and digital evolution,...
arXiv:2608. 03742v1 Announce Type: cross Abstract: Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability.
arXiv:2607. 03731v1 Announce Type: cross Abstract: Creating 3D assets for virtual reality requires modeling expertise, which restricts the authorship of immersive experiences.
arXiv:2607. 13560v1 Announce Type: cross Abstract: Recent advances in generative and embodied AI have been driven by large-scale predictive learning over multimodal data.
The article surveys how diffusion and flow-based generative models learn rich visual representations and how these representations can be used to improve generation and other perception tasks. It introduces a three-tier framework that categorizes work into improving generative quality via representation learning, extracting representations for perception, and developing unified applications. The survey covers downstream tasks such as image classification, dense prediction, instance-level perception, and annotation-scarce scenarios, offering a taxonomy and highlighting future research directions.
arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.