arXiv:2605.06272v2 Announce Type: replace
Abstract: While generative modeling has achieved remarkable success on tasks like natural language-conditioned image generation, enabling model adaptation fr...
By Tyler Ingebrand, Ruihan Zhao, Kushagra Gupta, David Fridovich-Keil, Sandeep P. Chinchali, Ufuk Topcu
arXiv:2610.00414v1 Announce Type: new
Abstract: Foundation models pretrained on large-scale datasets demonstrate strong transferability to medical imaging tasks. However, understanding how their late...
By Michael D. Vasilakakis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece), Dimitris K. Iakovidis (Department of Computer Science and Biomedical Informatics, University of Thessaly, Lamia, Greece)
arXiv:2610.01012v1 Announce Type: new
Abstract: Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental chal...
By Gunwoo Lee, Yoori Oh, Yoseob Han
arXiv:2610.01092v1 Announce Type: new
Abstract: Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must...
By Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji
arXiv:2610.01162v1 Announce Type: new
Abstract: Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, generated robot video...
By Isaiah Milkey, Som Sagar, Aditya Taparia, Xinyuan Liu, Jiqing Wen, Ransalu Senanayake
arXiv:2610.01233v1 Announce Type: new
Abstract: Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Repre...
By Zhen Zhou, Zhiwei Ning, Puhua Jiang, Sheng Zhang, Yifei Tang, Jie Yang, Xintong Han, Wei Liu, Chunchao Guo
arXiv:2610.01681v1 Announce Type: new
Abstract: Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-imag...
By Lidia Troeshestova, Alexander Ustyuzhanin, Sergey Kastryulin
arXiv:2610.01969v1 Announce Type: new
Abstract: Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-im...
By Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu
arXiv:2610.02044v1 Announce Type: new
Abstract: Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line...
By Tao Wu, Alexandra Gomez-Villa, Senmao Li, Yaxing Wang, Joost van de Weijer, Kai Wang
arXiv:2610.02153v1 Announce Type: new
Abstract: Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visu...
By Yiwen Zhang, Haocheng Xi, Michael Tian-Yue Liu, Alexei A. Efros, Hadar Averbuch-Elor, Qianqian Wang, Haiwen Feng
arXiv:2610.02208v1 Announce Type: new
Abstract: Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of...
By Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein
arXiv:2610.00895v1 Announce Type: cross
Abstract: Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such st...
By Sidney Bender, Benedikt Kunz, Ahmed Zeid, Shinichi Nakajima, Klaus-Robert M\"uller, Marco Morik
arXiv:2605.11506v2 Announce Type: replace
Abstract: Score-based diffusion models achieve state-of-the-art performance for inverse problems, but their practical deployment is hindered by long inferenc...
By Julio Oscanoa, Irmak Sivgin, Cagan Alkan, Daniel Ennis, John Pauly, Mert Pilanci, Shreyas Vasanawala
arXiv:2605.21244v2 Announce Type: replace
Abstract: Super-Resolution (SR) has advanced rapidly in recent years, with diffusion-based models achieving unprecedented fidelity at the cost of introducing...
By Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Dmitriy Vatolin
arXiv:2606.08091v2 Announce Type: replace
Abstract: Agentic long video generation requires planning, tool orchestration, and cross-clip coordination over a long horizon. Most existing video agents ei...
By Jianhui Wei, Yan Zhang, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Yeying Jin, Zuozhu Liu
arXiv:2607.08020v2 Announce Type: replace
Abstract: Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal co...
By Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le, Tam V. Nguyen, Minh-Triet Tran
arXiv:2610.01933v1 Announce Type: new
Abstract: Inference-time scaling adapts pretrained diffusion models to new sampling tasks without additional training. Existing methods rely primarily on Monte C...
By Zuokai Wen, Louis Grenioux, Weinan E, Jiequn Han
arXiv:2610.01522v1 Announce Type: cross
Abstract: Many scientific and machine learning systems, from molecular dynamics to diffusion models and beyond, are governed by stochastic dynamics with low-di...
By Vladimir R. Kostic, Karim Lounici, H\'el\`ene Halconruy, Timoth\'ee Devergne, Michele Parrinello, Massimiliano Pontil
arXiv:2610.02128v1 Announce Type: cross
Abstract: Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence struct...
By Shivam Kumar, Nabarun Deb
STAGE is a training‑free, closed‑form framework for concept erasure in native text‑to‑3D generators. It treats erasure as a stage‑aware editing problem, applying low‑dimensional affine corrections separately to the structural and appearance stages of the pipeline. Experiments on the TRELLIS generator show that STAGE outperforms adapted baselines, achieving a composite score of 66.7 versus 53.2 across 15 shape, material, and object concepts.
By Karol Dziekan, Przemys{\l}aw Spurek, Dawid Malarz