arXiv:2502. 06818v4 Announce Type: replace Abstract: Recent works modify CLIP to perform open-vocabulary semantic segmentation in a training-free manner (TF-OVSS).
By Jingyun Wang, Cilin Yan, Guoliang Kang
arXiv:2607. 15794v1 Announce Type: cross Abstract: Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion.
By Stefano Silvestrini, Michele Ceresoli
arXiv:2607. 15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning.
By Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
arXiv:2607. 15698v1 Announce Type: cross Abstract: We propose and evaluate three hierarchical ensemble setups for zebrafish phenotype classification from embryo images.
By Piotr S. Maci\k{a}g, Monika Maci\k{a}g, Magdalena Majdan
arXiv:2603. 02142v2 Announce Type: replace-cross Abstract: Scaling laws assume larger models trained on more data consistently outperform smaller ones -- an assumption that drives model selection in computer vision but remains untested in resource-constrained Earth observation (EO).
By Kwame Mbobda-Kuate, Gabriel Kasmi
arXiv:2607. 15477v1 Announce Type: new Abstract: Sleep apnea diagnosis via polysomnography remains resource intensive and relies on time consuming manual data analysis and scoring.
By Shashank Manjunath, Mukesh Cheemakurthi, Aarti Sathyanarayana
Hyperspectral image (HSI) classification systems are increasingly deployed on platforms with strict computational budgets, such as UAVs and small spaceborne sensors. In these settings, accuracy alone is not enough; the model must also run within tight latency and memory constraints.
arXiv:2408. 12548v3 Announce Type: replace Abstract: Machine Learning (ML) has become central to Autonomous Vehicles (AVs), supporting perception, prediction, planning, control, and decision-making in dynamic environments.
By Yousef Emami, Mohammadhossein Homaei, Miguel Guti\'errez Gait\'an, Luis Almeida, Kai Li, Hui Huang, Zhu Han
arXiv:2607. 14163v1 Announce Type: cross Abstract: Most single-cell foundation models are adapted from language models, representing each cell as a sequence of gene tokens.
By Ridvan Yesiloglu, Sakib Mostafa, James Zou, Ash Alizadeh, Jiajun Wu, Lei Xing, Ehsan Adeli, Md Tauhidul Islam
arXiv:2607. 14675v1 Announce Type: cross Abstract: Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources.
By Zihan Guo, Xiaoqi Li
arXiv:2604. 02429v2 Announce Type: replace-cross Abstract: Convolutional neural networks (CNNs) have transformed image processing, but the energy consumption and inference latency of electronic based implementations remain fundamental bottlenecks.
By Saurabh Ranjan, Sonika Thakral, Amit Sehgal
arXiv:2607. 14127v1 Announce Type: cross Abstract: Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant height of local obstructions that drive terminal clutter loss.
By Shohini Sarkar, Smithi Mahendran, Rishi Chudasama, Varun Mannam, Arav Luthra, Yuvraj Rekhi, Vivek Nadig, Arsh Goenka
arXiv:2607. 14703v1 Announce Type: cross Abstract: Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology.
By Mingxi Fu, Jiawen Li, Renao Yan, Jiali Hu, Qiehe Sun, Tian Guan, Yonghong He
arXiv:2607. 14287v1 Announce Type: cross Abstract: Defect segmentation in additive manufacturing (AM) X-ray computed tomography (XCT) images remains challenging due to severe class imbalance and large distribution shifts across scan conditions.
By Md Mahedi Hasan, Md Mushfiqur Rahaman, Alan Pachkovskiy, Imtiaz Ahmed, Jeremy Dawson, Srinjoy Das
arXiv:2607. 14160v1 Announce Type: new Abstract: Wildfire detection from satellite imagery is a semantic image segmentation problem that has proven to be difficult due to challenges such as class imbalance, feature complexity, and atmospheric interference.
By Jaiman Munshi (IonQ Team, App Dev Club, University of Maryland, College Park), Tanvi Tewary (IonQ Team, App Dev Club, University of Maryland, College Park), Sawyer Bloom (IonQ Team, App Dev Club, University of Maryland, College Park), Aidan Chu (IonQ Team, App Dev Club, University of Maryland, College Park), Chetan Maviti (IonQ Team, App Dev Club, University of Maryland, College Park), Kyon Winston-Bey (IonQ Team, App Dev Club, University of Maryland, College Park), Harshit Badjatia (IonQ Team, App Dev Club, University of Maryland, College Park), Farhan Kittur (IonQ Team, App Dev Club, University of Maryland, College Park), Vardhan Madhavarapu (IonQ Team, App Dev Club, University of Maryland, College Park), Varun Kota (IonQ Team, App Dev Club, University of Maryland, College Park), Joshua Kwon (IonQ Team, App Dev Club, University of Maryland, College Park), Nazia Rangwala-Vohra (IonQ Team, App Dev Club, University of Maryland, College Park), Franz Klein (IonQ Team, App Dev Club, University of Maryland, College Park)
arXiv:2607. 15082v1 Announce Type: cross Abstract: Understanding newspaper images remains a challenging task due to their complex, nested hierarchical structures and dense, heterogeneous layouts.
By William Moca\"er, Sol\`ene Tarride, Thomas Constum, Merveilles Agbeti-Messan, Tom Simon, Cl\'ement Chatelain, St\'ephane Nicolas, Pierrick Tranouez, S\'ebastien Cretin, Thierry Paquet
arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.
By Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin
arXiv:2607. 14443v1 Announce Type: new Abstract: Computer-use agents are becoming capable software operators, but their interface to desktop applications is still often a brittle motor layer: they look at screenshots, predict coordinates, click, and hope that the visible state changed as intended.
By Yong Liu, Zhenyi Zhong, Zhanpeng Shi
arXiv:2603. 16351v2 Announce Type: replace-cross Abstract: Accurate taxonomic identification of parasitoid wasps within the superfamily Ichneumonoidea is essential for biodiversity assessment, ecological monitoring, and biological control programs.
By Joao Manoel Herrera Pinheiro, Gabriela Do Nascimento Herrera, Alvaro Doria Dos Santos, Luciana Bueno Dos Reis Fernandes, Ricardo V. Godoy, Eduardo A. B. Almeida, Helena Carolina Onody, Marcelo Andrade Da Costa Vieira, Angelica Maria Penteado-Dias, Marcelo Becker
arXiv:2605. 25170v2 Announce Type: replace-cross Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models.
By Kordel K. France, Ovidiu Daescu