arXiv:2607. 23909v1 Announce Type: new Abstract: Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone.
By Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra
arXiv:2604. 22901v2 Announce Type: replace Abstract: Diffusion models achieve remarkable success in time series generation.
By Dong Liu, Yanxuan Yu, Ying Nian Wu
arXiv:2607. 23751v1 Announce Type: new Abstract: The usefulness of a variational autoencoder (VAE) depends on two properties of its latent space that are hard to obtain together: high encoding capacity in the individual latent variables, and a low-dimensional, disentangled organization of those variables.
By Ye Shi
arXiv:2607. 22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?
By Yassine Oueslati, Daniil Kirilenko, Martin Gjoreski, Marc Langheinrich
arXiv:2607. 22655v1 Announce Type: new Abstract: Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience.
By Jie Zhao, Jie Feng, Can Rong, Zhihan Hou, Peng Lu, Yong Li
arXiv:2603. 28963v2 Announce Type: replace-cross Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems.
By Mozhgan Pourkeshavarz, Tianran Liu, Nicholas Rhinehart
arXiv:2607. 22663v1 Announce Type: new Abstract: Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable.
By Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du
arXiv:2607. 24377v1 Announce Type: cross Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models.
By Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding, Yiwu Yao, Junsong Wang
arXiv:2607. 24741v1 Announce Type: cross Abstract: Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration.
By Xinyang Wen
arXiv:2509. 15357v3 Announce Type: replace-cross Abstract: Diffusion models have achieved strong results in text-to-image generation, but important limitations remain as prompts become more structured and multi-object.
By Yu Chang, Jiahao Chen, Anzhe Cheng, Paul Bogdan
arXiv:2607. 23946v1 Announce Type: new Abstract: We introduce Joint Flow Matching (JFM), a training framework for continuous normalising flows over multiple variables.
By Hayden McAlister, Lech Szymanski
arXiv:2607. 24176v1 Announce Type: cross Abstract: Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes.
By Alexey Gavrilov, Alan-Barsag Gazzaev, Sergey Muravyov
arXiv:2607. 23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation.
By Tanmay Khandait, Preetom Biswas, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Giulia Pedrielli
arXiv:2607. 24405v1 Announce Type: new Abstract: In this work, we propose K-SurvMeans, a novel extension of K-Means for clustering survival data.
By Abdallah Alabdallah
arXiv:2510. 03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastructure-based measures are often insufficient in dynamic urban environments.
By Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Muhammad Monjurul Karim, Kehua Chen, Chenxi Liu, Mehrdad Nasri, Yinhai Wang
arXiv:2604. 13472v2 Announce Type: replace-cross Abstract: Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a centralized control problem into multiple interacting agents.
By Zijian Zhao, Jing Gao, Sen Li
arXiv:2607. 24272v1 Announce Type: cross Abstract: The vast chemical design space and complex, interdependent design variables make catalyst discovery for targeted properties highly labor- and resource-intensive.
By Hayoung Doo, Dong Hyeon Mok, Seoin Back, Jonggeol Na
arXiv:2607. 24665v1 Announce Type: cross Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows.
By Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Xuelong Li
arXiv:2603. 25629v2 Announce Type: replace-cross Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs).
By Andr\'e G. Viveiros, Nuno Gon\c{c}alves, Matthias Lindemann, Andr\'e Martins
arXiv:2607. 23615v1 Announce Type: new Abstract: In multiple-input multiple-output (MIMO) semantic communication, imperfect channel state information (CSI) and equalization mismatch can seriously degrade semantic reconstruction quality.
By Wenkai Liu, Nan Ma, Jianqiao Chen, Xiaodong Xu, Meixia Tao, Ping Zhang