arXiv:2603. 17555v2 Announce Type: replace-cross Abstract: Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.
By Hugo Caselles-Dupr\'e, Mathis Koroglu, Guillaume Jeanneret, Arnaud Dapogny, Matthieu Cord
arXiv:2605. 10723v2 Announce Type: replace-cross Abstract: Generating a complete music video from a song requires more than synthesizing visually plausible clips for individual lyric prompts.
By Huimin Wang, Chang Xia, Leilei Ouyang, Yongqi Kang, Yu Fu, Yuqi Ouyang
arXiv:2608. 16513v1 Announce Type: cross Abstract: Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation.
By Junhao Chen, Zheqi Lv, Keting Yin, Shengyu Zhang, Zhou Zhao, Feiyang Chen, Xinyu Duan, Baoxing Huai, Fei Wu
arXiv:2608. 16380v1 Announce Type: cross Abstract: Monitoring war-induced damage to agricultural land in Ukraine is important for understanding threats to food security, environmental stability, and post-war recovery.
By Marta Sumyk, Oleksandr Kosovan, Iryna Voitsitska
arXiv:2608. 15694v1 Announce Type: cross Abstract: Conditional image-to-image generators are single-shot: they map input features to an output in one forward pass and treat it as final, with no opportunity to improve on it.
By Kareem Hassani, Chaymaa Abbas, Hadi Al Mubasher, Mariette Awad
arXiv:2608. 15881v1 Announce Type: new Abstract: The Multiphysics Object-Oriented Simulation Environment (MOOSE) is an open-source finite-element framework for building multiphysics simulation applications.
By Zaid Abulawi, Mengnan Li, Guillaume Giudicelli, Yang Liu, Cody Permann
arXiv:2608. 15768v1 Announce Type: cross Abstract: Conformal prediction (CP) provides distribution-free coverage guarantees and has emerged as a principled tool for uncertainty quantification.
By Xudong Chen, Shengbo Gong, Lu Cheng, Wei Jin
arXiv:2608. 15037v1 Announce Type: cross Abstract: Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference.
By Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
arXiv:2608. 16473v1 Announce Type: new Abstract: Neural PDE training yields a finite checkpoint archive, yet its logged energy errors are inaccessible without the exact solution, while loss-based selection does not necessarily recover the logged energy oracle.
By Karim Bounja, Lahcen Laayouni, Boujemaa Achchab, Abdeljalil Sakat
arXiv:2608. 16569v1 Announce Type: new Abstract: Accurate reconstruction of long-duration neural recordings is challenging because local field potentials (LFPs) are high-resolution, multichannel, transient, and variable across subjects.
By Anima Kujur, Zahra Monfared
arXiv:2608. 16765v1 Announce Type: cross Abstract: Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.
By Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
arXiv:2509. 24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data.
By Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang
arXiv:2608. 16457v1 Announce Type: cross Abstract: Procedure planning seeks to estimate a sequence of actions to transition from an observed initial state to a given goal state.
By Mohamed Afham, Christoph Reich, Oliver Hahn, Daniel Cremers, Stefan Roth
arXiv:2608. 14758v1 Announce Type: cross Abstract: For patients with Moyamoya disease, impaired cerebrovascular reserve (CVR) is an important hemodynamic criterion for recommending extracranial-to-intracranial bypass surgery.
By Julia Huang, Camila Gonzalez, Rydham Goyal, Aja Zou, Sasha Alexander, Michael Moseley, Moss Y. Zhao, Gary K. Steinberg
arXiv:2608. 15958v1 Announce Type: new Abstract: Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no short certificate to check.
By Sina Baghal
arXiv:2608. 16791v1 Announce Type: cross Abstract: Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities.
By Ye Lu, Shen Wang, Zhaoyang Zhang, Yihan Yan, Li Liu, Runze Liu, Fanghui Sun
arXiv:2608. 15314v1 Announce Type: new Abstract: Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) falls below an outage threshold.
By Amanda Sheron Gamage, Niloofar Mehrnia, James Gross
arXiv:2608. 14599v1 Announce Type: cross Abstract: The placement of base station (BS) is a fundamental determinant of coverage and capacity of urban wireless networks.
By Zhenyu Tao, Yuxuan Li, Wei Xu, Yongming Huang, Xiaohu You
arXiv:2608. 14649v1 Announce Type: new Abstract: We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification.
By Pawan Kumar
arXiv:2602. 22431v2 Announce Type: replace-cross Abstract: Millimeter-wave (mmWave) radar captures are band-limited and noisy, making for difficult reconstruction of intelligible full-bandwidth speech.
By Jash Karani, Adithya Chittem, Deepan Roy, Sandeep Joshi