arXiv:2607. 13624v1 Announce Type: cross Abstract: Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms.
By Jose Mart\'inez-Fajardo, Pablo Pueyo, Fernando Caballero, Luis Merino
arXiv:2607. 13394v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes.
By Xiaodong Liu, Michael Xu, Jack W. Stokes, Paul Smolensky, Doug Burger, Jianfeng Gao
arXiv:2606. 05981v2 Announce Type: replace-cross Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1-step distilled student, the text encoder becomes the critical path.
By Yoshiyuki Ootani
arXiv:2603. 14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models.
By Charlotte Morissette, Amin Abyaneh, Wei-Di Chang, Anas Houssaini, David Meger, Hsiu-Chin Lin, Jonathan Tremblay, Gregory Dudek
arXiv:2607. 13597v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization.
By Yuan Xu, Youheng Shi, Chengyang Li, Wentao Zhu, Yizhou Wang
arXiv:2607. 13188v1 Announce Type: new Abstract: Human cognition does not separate understanding and generation.
By Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro V\'elez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
arXiv:2603. 00546v2 Announce Type: replace Abstract: Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains.
By Zeyu Chen, Huanjin Yao, Ziwang Zhao, Min Yang
arXiv:2607. 13395v1 Announce Type: new Abstract: The pursuit of autonomously self-improving models has attracted growing interest in the era of large-scale foundation models.
By Jing-Xiao Liao, Tianwei Zhang, Yu-Hao Jiang, Feifei Zhang, Hang-Cheng Dong, Feng-Lei Fan
arXiv:2510. 04876v3 Announce Type: replace-cross Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource management.
By Hayat Rajani, Valerio Franchi, Borja Martinez-Clavel Valles, Raimon Ramos, Rafael Garcia, Nuno Gracias
arXiv:2604. 06614v2 Announce Type: replace-cross Abstract: Prompt learning has gained significant attention as a parameter-efficient approach for adapting large pre-trained vision-language models to downstream tasks.
By Yaqi Zhao, Haoliang Sun, Yating Wang, Yongshun Gong, Yilong Yin
arXiv:2508. 16403v3 Announce Type: replace Abstract: Accurately predicting the performance of active radio frequency (RF) circuits is essential for modern wireless systems but remains challenging due to highly nonlinear behavior and the high computational cost of traditional simulation tools.
By Anahita Asadi, Leonid Popryho, Inna Partin-Vaisband
arXiv:2605. 31272v2 Announce Type: replace Abstract: As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provide recourse to affected individuals.
By Wenshuo Dong, Jiaming Zhang, Shaopeng Fu, Hongbin Lin, Di Wang, Lijie Hu
arXiv:2607. 13984v1 Announce Type: cross Abstract: Longitudinal tumor measurements, dropout information, and genetic covariates provide complementary information about treatment response, but integrating these data sources within a single population modeling framework remains challenging.
By Anders Sj\"oberg, Nils Olsson, Marcus Baaz, Mats Jirstrand
arXiv:2508. 12466v2 Announce Type: replace-cross Abstract: Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting visual features into discrete text token spaces using large-scale image--text data.
By Xuhui Zhan, Tyler Derr
arXiv:2607. 13110v1 Announce Type: cross Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences.
By Yi Wang, Yinfeng Yu
arXiv:2607. 13239v1 Announce Type: new Abstract: Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center (TMC) tasks such as anomaly detection, incident reporting, and traveler information.
By Xi Cheng, Ke Liu, Siyuan Feng, Jane Lin, H. Oliver Gao
arXiv:2607. 13881v1 Announce Type: cross Abstract: Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories.
By Ting Lei, Jialin Liu, Zhu Xu, Yuxin Peng, Yang Liu
arXiv:2607. 13164v1 Announce Type: cross Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly because video diffusion models are expensive to train and evaluate.
By Ruize Xia
arXiv:2607. 13548v1 Announce Type: new Abstract: Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches.
By Athira Gopal, Ashwanth Krishnan
arXiv:2607. 13826v1 Announce Type: cross Abstract: Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability.
By Vincent Ochs, Christoph Kuemmerli, Florentin Bieder, Julia Wolleb, Joel L. Lavanchy, Julia Ruppel, Jan Liechti, Stephanie Taha-Mehlitz, Christian Andreas Nebiker, Beat Mueller, Giuseppe Kito Fusai, Joerg-Matthias Pollok, Anas Taha, Philippe C. Cattin, Sebastian Staubli