arXiv:2609.22688v1 Announce Type: new
Abstract: Generating parametric CAD models requires accurate geometry and stable feature dependencies. Existing methods face challenges in selecting geometric re...
By Xi Cheng, Chenxi Zhai, Hang Cheng, Mingyu Fan, Pingfa Feng, Long Zeng
arXiv:2609.22789v1 Announce Type: new
Abstract: Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusi...
By Zelin Jia, Zhao Zhang, Zhicong Tang, Yuhui Yuan, Shixia Liu
arXiv:2609.22834v1 Announce Type: new
Abstract: Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation...
By Changhao Zhao, Linglin Zeng, Hai Liu
arXiv:2609.23012v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene...
By Moshiur Farazi, Bekir Ciftler, Abdulhalim Dandoush, Reda Bendraou
arXiv:2609.23121v1 Announce Type: new
Abstract: Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical framework...
By Yang Tian, Fan Liu, Jingyuan Zhang, Zhenyang Li, Yupeng Hu, Liqiang Nie
arXiv:2609.23248v1 Announce Type: new
Abstract: A fundamental challenge in deploying vision models is domain shift, which arises when training and test data follow different distributions, leading to...
By Jo\~ao Renato Ribeiro Manesco, Danilo Samuel Jodas, Douglas Rodrigues, Leandro Aparecido Passos, Jo\~ao Paulo Papa
arXiv:2609.23345v1 Announce Type: new
Abstract: As generative AI becomes increasingly used in anime-style image creation, distinguishing human-drawn, AI-inpainted, and text-to-image images is importa...
By Yan Liu, Baoxiang Huang, Zi'an Wang, Wenbo Xie
arXiv:2609.23541v1 Announce Type: new
Abstract: Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and...
By Ziying Song, Lin Liu, Hongyu Pan, Shaoqing Xu, Lei Yang, Mingzhe Guo, Caiyan Jia
arXiv:2609.23565v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot contr...
By Yuxuan Jiang, Jiaying Huang, Ge Wang, Shenhao Yan, Jiahao Yang, Chengsi Yao, Qi Liu, Qing Zhao, Shuguang Cui, Yiming Zhao, Yatong Han, Zhen Li
arXiv:2609.23655v1 Announce Type: new
Abstract: Imbalanced gradient magnitudes between the visual and language pathways of a vision-language model are often treated as a defect to be corrected. We te...
By Kiran Naseer, Samreen Azhar, Dwarikanath Mahapatra
arXiv:2609.23769v1 Announce Type: new
Abstract: The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-co...
By Manuel Serna-Aguilera, Raegan Anderes, Page Dobbs, Khoa Luu
arXiv:2609.23796v1 Announce Type: new
Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challe...
By Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi
arXiv:2609.24190v1 Announce Type: new
Abstract: Artificial intelligence (AI) has advanced at a rapid pace in recent years. Initially, breakthroughs in large language models caught widespread attentio...
By Rian Dolphin, Laura Knowles
arXiv:2609.24470v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly explored as interfaces for scientific image analysis, where a visual question-answering (VQA)...
By Samia Mohinta, Albert Cardona
arXiv:2609.24526v1 Announce Type: new
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and...
By Foundation Model, Li Auto Inc
arXiv:2609.24539v1 Announce Type: new
Abstract: Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic...
By Weixiang Zhou, Yuhao Wang, Xingguo Xu, Weizhen Zhou, Zhixun Su, Jinshan Pan, Cong Wang
arXiv:2609.24564v1 Announce Type: new
Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encode...
By Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen
arXiv:2609.22277v1 Announce Type: cross
Abstract: Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environm...
By Ali Akarma, Adeel Ahmad, Toqeer Ali Syed
arXiv:2609.23417v1 Announce Type: cross
Abstract: Screen-recorded web editing requests contain weak deictic expressions such as ``this'' and ``there,'' whose referents depend on speech, cursor trajec...
By Minghao Han, Zhenghao Xing, Xize Cheng, Yuxuan Wang, Junming Lin, Ling Wang, Yinsong Yan, Yunfei Chu, Qize Yang, Jin Xu
arXiv:2609.24187v1 Announce Type: cross
Abstract: Autonomous endoscopic navigation requires the policy model to predict actions from texture-poor monocular observations, make safe control decisions,...
By Tamima Tabassum, Yiming Huang, Tianchun Wu, Changjing Liu, Zhiqing Tang, Chikit Ng, Beilei Cui, Liangjing Shao, Jiewen Lai, Hongliang Ren