Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

1,466 stories · RSS feed

arXiv AI
Jul 20

On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation

arXiv:2607. 15794v1 Announce Type: cross Abstract: Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion.

By Stefano Silvestrini, Michele Ceresoli
arXiv Machine Learning
Jul 17

Human-In-The-Loop Machine Learning for Safe and Ethical Autonomous Vehicles: Principles, Challenges, and Opportunities

arXiv:2408. 12548v3 Announce Type: replace Abstract: Machine Learning (ML) has become central to Autonomous Vehicles (AVs), supporting perception, prediction, planning, control, and decision-making in dynamic environments.

By Yousef Emami, Mohammadhossein Homaei, Miguel Guti\'errez Gait\'an, Luis Almeida, Kai Li, Hui Huang, Zhu Han
arXiv AI
Jul 17

Explainable Geospatial AI for Satellite Ground Station Siting Using LiDAR-Derived Terrain Intelligence

arXiv:2607. 14127v1 Announce Type: cross Abstract: Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant height of local obstructions that drive terminal clutter loss.

By Shohini Sarkar, Smithi Mahendran, Rishi Chudasama, Varun Mannam, Arav Luthra, Yuvraj Rekhi, Vivek Nadig, Arsh Goenka
arXiv Machine Learning
Jul 17

XCT-SAM: Sequential Parameter-Efficient Domain Adaptation of SAM for Industrial XCT Defect Segmentation

arXiv:2607. 14287v1 Announce Type: cross Abstract: Defect segmentation in additive manufacturing (AM) X-ray computed tomography (XCT) images remains challenging due to severe class imbalance and large distribution shifts across scan conditions.

By Md Mahedi Hasan, Md Mushfiqur Rahaman, Alan Pachkovskiy, Imtiaz Ahmed, Jeremy Dawson, Srinjoy Das
arXiv Machine Learning
Jul 17

QFireNet: A Quantum-Enhanced U-Net for Wildfire Segmentation from Sentinel-2 Imagery

arXiv:2607. 14160v1 Announce Type: new Abstract: Wildfire detection from satellite imagery is a semantic image segmentation problem that has proven to be difficult due to challenges such as class imbalance, feature complexity, and atmospheric interference.

By Jaiman Munshi (IonQ Team, App Dev Club, University of Maryland, College Park), Tanvi Tewary (IonQ Team, App Dev Club, University of Maryland, College Park), Sawyer Bloom (IonQ Team, App Dev Club, University of Maryland, College Park), Aidan Chu (IonQ Team, App Dev Club, University of Maryland, College Park), Chetan Maviti (IonQ Team, App Dev Club, University of Maryland, College Park), Kyon Winston-Bey (IonQ Team, App Dev Club, University of Maryland, College Park), Harshit Badjatia (IonQ Team, App Dev Club, University of Maryland, College Park), Farhan Kittur (IonQ Team, App Dev Club, University of Maryland, College Park), Vardhan Madhavarapu (IonQ Team, App Dev Club, University of Maryland, College Park), Varun Kota (IonQ Team, App Dev Club, University of Maryland, College Park), Joshua Kwon (IonQ Team, App Dev Club, University of Maryland, College Park), Nazia Rangwala-Vohra (IonQ Team, App Dev Club, University of Maryland, College Park), Franz Klein (IonQ Team, App Dev Club, University of Maryland, College Park)
arXiv AI
Jul 17

Towards Hierarchical Structure Understanding of Newspaper Images

arXiv:2607. 15082v1 Announce Type: cross Abstract: Understanding newspaper images remains a challenging task due to their complex, nested hierarchical structures and dense, heterogeneous layouts.

By William Moca\"er, Sol\`ene Tarride, Thomas Constum, Merveilles Agbeti-Messan, Tom Simon, Cl\'ement Chatelain, St\'ephane Nicolas, Pierrick Tranouez, S\'ebastien Cretin, Thierry Paquet
arXiv AI
Jul 17

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

arXiv:2607. 14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time.

By Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin
arXiv AI
Jul 17

Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI

arXiv:2603. 16351v2 Announce Type: replace-cross Abstract: Accurate taxonomic identification of parasitoid wasps within the superfamily Ichneumonoidea is essential for biodiversity assessment, ecological monitoring, and biological control programs.

By Joao Manoel Herrera Pinheiro, Gabriela Do Nascimento Herrera, Alvaro Doria Dos Santos, Luciana Bueno Dos Reis Fernandes, Ricardo V. Godoy, Eduardo A. B. Almeida, Helena Carolina Onody, Marcelo Andrade Da Costa Vieira, Angelica Maria Penteado-Dias, Marcelo Becker