arXiv:2609.00901v1 Announce Type: new
Abstract: Modifying the illumination of driving images is a fundamental challenge, as most datasets are captured at specific times of day. Existing methods rely...
By Hala Djeghim, Nathan Piasco, Luis Rold\~ao, Moussab Bennehar, Dzmitry Tsishkou, C\'eline Loscos, D\'esir\'e Sidib\'e
arXiv:2607. 01983v1 Announce Type: cross Abstract: Robust 3D object detection under adverse weather remains a critical hurdle for autonomous driving.
By Shuyao Li, Chuanxing Geng, Heyang Sun, Qiang Zhou, Jingjing Gu
Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion to the planning module, conditioning on fixed outputs from separate discriminative perception networks.
arXiv:2609.36810v1 Announce Type: new
Abstract: Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determi...
By Renlong Wu, Guanqiao Wang, Xuan Shang, Yin Hanming, Xiaoxiao Sheng, Tianyu Huang, Hui Li, Wangmeng Zuo
arXiv:2606. 29020v1 Announce Type: cross Abstract: Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion.
By Chenghao Qian, Nedko Savov, Lingdong Kong, Yeying Jin, Rui Song, Wenjing Li, Zhun Zhong, Jiaqi Ma, Gustav Markkula, Luc Van Gool
Bernini proposes a unified framework that separates semantic planning and pixel rendering for video generation and editing. An MLLM-based planner predicts target semantics in ViT embedding space, while a DiT-based renderer synthesizes pixels conditioned on this plan, text features, and source VAE features for editing. The approach introduces Segment-Aware 3D Rotary Positional Embedding and chain-of-thought reasoning, achieving state‑of‑the‑art performance on diverse video benchmarks.
By Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input.
3D semantic scene generation is crucial for autonomous driving applications, yet most methods rely on complex 3D-specific architectures such as triplane encoders and adapted diffusion networks, limiting both their simplicity and their editing capabilities. We propose EditSSC, an editing-ready method for 3D semantic scene generation using 2D Bird's Eye View (BEV) representations and off-the-shelf latent diffusion network.
arXiv:2411.05005v2 Announce Type: replace-cross
Abstract: Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, m...
By Shuhong Zheng, Zhipeng Bao, Ruoyu Zhao, Martial Hebert, Yu-Xiong Wang
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
By Habin Lim, Gyeong-Moon Park
arXiv:2607. 09764v1 Announce Type: cross Abstract: The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust autonomous driving systems.
By Xiaoyun Dong, Qian Xu, Yang Lu, Yang Lou, Yung-Hui Li, Jianping Wang
arXiv:2605.12957v2 Announce Type: replace
Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of do...
By Hanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo, Tianyu He, Xin Jin, Zhibo Chen