arXiv:2609.36429v1 Announce Type: new
Abstract: Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods...
By Zijun Gao, Chunbin Gu, Jinxi Xiang, Xiangde Luo, Pheng-Ann Heng
arXiv:2609.36598v1 Announce Type: new
Abstract: A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is un...
By Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng, Siyu Zhu, Long Qin, Zhenghao Zhang
arXiv:2609.36826v1 Announce Type: new
Abstract: Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-...
By Zhenghao Ni, Weimin Qiu, Meng Tang
arXiv:2609.36844v1 Announce Type: new
Abstract: Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the b...
By Suhani Grover, Astik Srivastava, Viswas Dinesh, Avinash Sharma, K. Madhava Krishna
arXiv:2609.36918v1 Announce Type: new
Abstract: Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due...
By Jiantao Lin, Meixi Chen, Yingjie Xu, Chenbo Fu, Leyi Wu, Hao Chen, Yinchuan Li, Ying-Cong Chen
arXiv:2609.36929v1 Announce Type: new
Abstract: Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However,...
By Thai Duy Nguyen, Addison Lin Wang
arXiv:2609.37003v1 Announce Type: new
Abstract: Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most e...
By Danfeng Hong, Chenyu Li, Jocelyn Chanussot
arXiv:2609.37015v1 Announce Type: new
Abstract: We propose RBF-GNN, a new pseudo-coordinate based graph neural network architecture that takes into account Euclidean, spherical or angular coordinates...
By Pawe{\l} Batorski, Abtin Pourhadi, Paul Swoboda
arXiv:2609.37141v1 Announce Type: new
Abstract: Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Exi...
By Xinye Yang, Xinding Zhu, Kai Fang, Xinyi Ren, Mengjian Li, Bin Cao, Jiazhou Chen
arXiv:2609.37283v1 Announce Type: new
Abstract: Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evide...
By Xuyang Cao, Enyou Liu, Jun Zhao, Zhuoyun Liu, Jintao Fei, Leo
arXiv:2609.37339v1 Announce Type: new
Abstract: General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on boundin...
By Jer Pelhan, Alan Lukezic, Matej Kristan
arXiv:2609.37870v1 Announce Type: new
Abstract: Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomo...
By Zhixiang Hao, Shaodi You, Yu Li, Kunming Li, Feng Lu
arXiv:2609.38079v1 Announce Type: new
Abstract: Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and obje...
By Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
arXiv:2609.37602v1 Announce Type: cross
Abstract: Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where envi...
By Michele Antonazzi, Alejandra C. Hernandez, Jos\'e Araujo, Olov Andersson, Patric Jensfelt
arXiv:2411.02562v2 Announce Type: replace
Abstract: Segmentation of cellular structures in electron microscopy (EM) images is fundamental to analyzing the morphology of neurons and glial cells in the...
By Zewen Zhuo, Ilya Belevich, Ville Leinonen, Eija Jokitalo, Tarja Malm, Alejandra Sierra, Jussi Tohka
arXiv:2507.13628v3 Announce Type: replace
Abstract: Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understand...
By Masahiro Ogawa, Qi An, Atsushi Yamashita
arXiv:2508.17171v3 Announce Type: replace
Abstract: Imaging biomarkers in magnetic resonance imaging (MRI) are important tools for diagnosing, tracking and treating Alzheimer's disease (AD). Neurofib...
By Yue Li, Pulkit Khandelwal, Rohit Jena, Long Xie, Michael Duong, Amanda E. Denning, Christopher A. Brown, Laura E. M. Wisse, Sandhitsu R. Das, David A. Wolk, Paul A. Yushkevich
arXiv:2512.17323v2 Announce Type: replace
Abstract: Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video f...
By Jiyun Kong, Jun-Hyuk Kim, Jong-Seok Lee
arXiv:2603.14733v2 Announce Type: replace
Abstract: Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos...
By Yue Zhang, Liqiang Jing, Jia Li, Yapeng Tian, Xinya Du, Yunhui Guo, Vibhav Gogate
arXiv:2603.12789v3 Announce Type: replace
Abstract: Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independ...
By Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park