arXiv:2608. 03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts.
By Ning Zhu, An Chen, Mengfei Zhao, Juntao Xu, Jingze Liang, Boyuan Gu, Liang-Jian Deng
arXiv:2608. 03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve.
By Fan Yang, Yuting Su, Xiaobo Wang, Yuncheng You, Fugui Fan, Yuting Wu, Minghui Wu, Chenxu Zhao, JiaHong Ning, Peiguang Jing
arXiv:2604. 18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adaptation based on the outcomes of multiple past trials.
By Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song
arXiv:2608. 03103v1 Announce Type: cross Abstract: Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation.
By Rishabh Shukla, Adithya Santhosh, Shaili Gandhi, Samrudh Moode, Satyandra K. Gupta
arXiv:2608. 03469v1 Announce Type: cross Abstract: We study score learning for reflected diffusion on bounded domains.
By Ziyue Wang, Takafumi Kanamori
arXiv:2608. 03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped.
By Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
arXiv:2608. 03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility.
By Mingzhuo Li, Guang Li, Linfeng Ye, Jiafeng Mao, Takahiro Ogawa, Konstantinos N. Plataniotis, Miki Haseyama
arXiv:2608. 03822v1 Announce Type: cross Abstract: Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation.
By Zhang Weihui, Wang Ruizhi, Xu Hongye, Wang Huiqiong, Sun Li, Song Mingli
arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.
By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv:2608. 03878v1 Announce Type: new Abstract: Synthetic power-grid scenarios are essential for planning, resilience assessment, contingency analysis, and data-driven power-system applications.
By Chenhan Xiao, Xinyu He, Haoran Li, Hanghang Tong, Yang Weng
arXiv:2608. 03413v1 Announce Type: new Abstract: As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image generation tasks, increasingly integrating tools, agents, and harnesses to solve real business and industrial problems.
By Zuojun Max Shen, Yuan Qu, Pujun Zhang, Anbang Liu, Yunhao Liang
arXiv:2608. 03260v1 Announce Type: new Abstract: Pretraining has shown strong potential for learning transferable representations, yet it remains underexplored for electron-density-based molecular learning.
By Liang Shuang, Haocheng Wang, Jiayi Song, Shuquan Ye, Ben Fei
arXiv:2607. 27933v2 Announce Type: replace Abstract: Flow matching (FM) has become a popular action head paradigm for modern embodied models.
By Ziyang Rao, Yiren Zhao, Weiyu Guo, Ben Fei, Yandong Guo, Hui Xiong
arXiv:2608. 03284v1 Announce Type: cross Abstract: Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can often elicit prohibited content, e.
By Jinya Sakurai, Shueicheng Yan, Xun Xu
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time, open-ended video editing without access to future frames or a predefined video duration.
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generation. We prove unconditional separations between low-depth quantum computation and the corresponding bounded-resource classical language-model architectures in both regimes.
Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation. Although generative models provide a promising means of data augmentation, existing methods often yield implausible spatial layouts of flooded regions and distort scene structures.
Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under different frames of reference.
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts.
Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing methods target randomly initialized networks, whereas modern vision systems often adapt frozen pretrained encoders with lightweight modules.