arXiv:2607. 23384v1 Announce Type: cross Abstract: Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy depends critically on associating measurements with the correct landmark variables.
By Yihao Zhang, Jungseok Hong, John J. Leonard
arXiv:2607. 22959v1 Announce Type: cross Abstract: AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control remains a major constraint to scaling production.
By Aniket Sakpal, Yang Jiang, Rouzbeh Davoudi, Shayan Hassantabar, Mani Najmabadi
arXiv:2607. 22743v1 Announce Type: cross Abstract: Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion.
By Madan Baduwal, Priyanka Paudel
arXiv:2607. 24512v1 Announce Type: new Abstract: Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles.
By Jan Range, Bj\"orn Schembera, Dominik G\"oddeke
arXiv:2607. 22611v1 Announce Type: new Abstract: The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access control (RBAC) models cannot address.
By Arun Malik, Deepal Jayasinghe, Bradley Klemick, Prachi Shah, Nitish Talasu, Vineet Tushar Trivedi
arXiv:2607. 23147v1 Announce Type: cross Abstract: Large language models now power autonomous agents capable of complex, multi-step tasks in different environments.
By Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck
arXiv:2607. 23245v1 Announce Type: cross Abstract: Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients.
By Haochen Liang, Jie Zhang, Hideya Ochiai
arXiv:2607. 23493v1 Announce Type: cross Abstract: Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age.
By Musa Tur Farazi, Nufayer Jahan Reza
arXiv:2607. 23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions.
By Eli Levinkopf, Efrat Morin, Claudia V. Goldman
arXiv:2607. 22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world.
By Wang Yang, Shaobo Wang, Hongxuan Liu, Xiaoran Cai, Yunyu He, Jingzong Zhou, Mengzhong Ma, Yi Yu, Rohit Sharma, Jingjing Fu, Peng Qi
arXiv:2607. 22745v1 Announce Type: cross Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation.
By Yi-Zhi Wang, Yichen Xiao, Linan Yue, Weibo Gao, Yichao Du, Pengfei Fang, Shimin Di, Min-Ling Zhang
arXiv:2607. 24556v1 Announce Type: cross Abstract: Split learning enables collaborative model training by partitioning neural networks across clients and servers.
By Akarsh K. Nair, Muhammad Arifur Rahman, David Brown, Mufti Mahmud
arXiv:2607. 22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior.
By Shawn Ray
arXiv:2607. 22702v1 Announce Type: cross Abstract: Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI.
By Addison Zucek, Prerit Gupta, Kamila Kuatova, Aniket Bera
arXiv:2607. 22985v1 Announce Type: cross Abstract: Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models.
By Chengyao Yu, Hongxin Wei, Bingyi Jing
arXiv:2607. 24683v1 Announce Type: cross Abstract: Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance.
By Francisco Mena, Dino Ienco, Roberto Interdonato, Cassio F. Dantas, Simon Besnard
arXiv:2604. 11912v2 Announce Type: replace-cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks.
By Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
arXiv:2606. 19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information.
By Yijin Wang, Shuyi Wang, Wenhan Zhang, Yuqi Ouyang
arXiv:2607. 22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs).
By Jian Hong, Chen Cheng, Quan Liu, Yuhao Chen, Enhong Chen
arXiv:2607. 22568v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency.
By Ruiyi Tao, Xiaolong Tu, Haoxin Wang