arXiv AI

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

Diffusion TV is an interactive AI art installation that transforms a CRT TV into a tangible interface for diffusion models. By physically adjusting the TV’s antenna, users control the clarity of AI-generated images and sounds, symbolically mirroring the denoising process that underlies diffusion-based generation. The installation offers three channels—Past, Present, and Future—featuring AI-generated animals, allowing participants to explore intermediate states of the generative process through continuous audiovisual feedback.

arXiv AI
Sep 24

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

Passing is an interactive audiovisual installation that transforms a single continuous monorail-window recording into an endless journey by reconstructing it as a spatiotemporal volume and resampling its spatial and temporal structure along nonlinear trajectories. A camera-based viewer‑presence detection system influences transitions among rendered video sequences, and the resulting video stream is fed into SpecMaskFoley, a real‑time video‑to‑audio synthesis model that generates a synchronized soundscape. The work distributes creative agency among the artist, the AI model, and the audience, exploring how authorship and listening can be negotiated among human intention, machine inference, and audience interpretation.

By Akira Takahashi, Chihiro Nagashima, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji
arXiv Computer Vision
Aug 26

Representation Learning in Diffusion and Flow-based Model: An Application Aspect

The article surveys how diffusion and flow-based generative models learn rich visual representations and how these representations can be used to improve generation and other perception tasks. It introduces a three-tier framework that categorizes work into improving generative quality via representation learning, extracting representations for perception, and developing unified applications. The survey covers downstream tasks such as image classification, dense prediction, instance-level perception, and annotation-scarce scenarios, offering a taxonomy and highlighting future research directions.

By Yanchen Xu, Sida Huang, Zhenyu Gu, Ruishu Zhu, Yilan Gao, Hongyuan Zhang
arXiv AI
Jul 8

Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

arXiv:2605. 13974v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood.

By Evelyn Turri, Davide Bucciarelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia