Diffusion and generative media

Image, video and audio generation — diffusion models, flow matching and the systems built on top of them.

2,793 stories · RSS feed

arXiv Machine Learning
Aug 7

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

arXiv:2608. 05728v1 Announce Type: cross Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval.

By Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan
arXiv AI
Aug 7

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

arXiv:2608. 05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emph{Der Giftpilz}.

By Ye Leng, Junjie Chu, Yiting Qu, Mingjie Li, Yun Shen, Yang Zhang
arXiv Machine Learning
Aug 7

A Reverse-BSDE Diffusion Sampler

arXiv:2505. 06800v2 Announce Type: replace-cross Abstract: Diffusion-based generative models have renewed interest in stochastic differential equation methods for sampling from complex distributions.

By Jairon H. N. Batista, Fl\'avio B. Gon\c{c}alves, Yuri F. Saporito, Rodrigo S. Targino
arXiv Machine Learning
Aug 7

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

arXiv:2608. 06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models.

By Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet, Romuald Elie
arXiv AI
Aug 7

SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

arXiv:2608. 05970v1 Announce Type: cross Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks.

By Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu
arXiv AI
Aug 7

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

arXiv:2608. 06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age.

By Arya Labroo, Mengjie Qian, Kate Knill
arXiv Machine Learning
Aug 7

KVAE: Family of Tokenizers for Multimodal Generative Models

arXiv:2608. 05798v1 Announce Type: cross Abstract: Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation.

By Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov, Kirill Chernyshev, Kirill Malakhov, Ilia Vasiliev, Ilia Trushkin, Valeriya Kobenko, David Chikovani, Alexander Ivanov, Azat Saginbaev, Egor Silvestrov, Ivan Mikheev, Konstantin Zakharov
arXiv Machine Learning
Aug 7

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

arXiv:2608. 06259v1 Announce Type: new Abstract: Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations.

By Yiting Zheng, Cheng Fang, Anthony Donofrio, Haote Li