Diffusion and generative media

Image, video and audio generation — diffusion models, flow matching and the systems built on top of them.

5,061 stories · RSS feed

arXiv Computer Vision
1d ago

From Image Latent Space to Fuzzy Rules: Interpretable Analysis of Gastrointestinal Foundation Model

arXiv:2610.00414v1 Announce Type: new Abstract: Foundation models pretrained on large-scale datasets demonstrate strong transferability to medical imaging tasks. However, understanding how their late...

By Michael D. Vasilakakis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece), Dimitris K. Iakovidis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
arXiv Computer Vision
1d ago

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

arXiv:2610.01092v1 Announce Type: new Abstract: Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must...

By Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji
arXiv Computer Vision
1d ago

Sphere Encoder 2

arXiv:2610.02208v1 Announce Type: new Abstract: Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of...

By Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein
arXiv Computer Vision
1d ago

STAGE: Subspace-Targeted Affine Generative Erasure for Text-to-3D Models

STAGE is a training‑free, closed‑form framework for concept erasure in native text‑to‑3D generators. It treats erasure as a stage‑aware editing problem, applying low‑dimensional affine corrections separately to the structural and appearance stages of the pipeline. Experiments on the TRELLIS generator show that STAGE outperforms adapted baselines, achieving a composite score of 66.7 versus 53.2 across 15 shape, material, and object concepts.

By Karol Dziekan, Przemys{\l}aw Spurek, Dawid Malarz