arXiv:2307. 00919v2 Announce Type: replace-cross Abstract: Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification benchmarks.
By Vinoth Nandakumar, Arush Tagade, Tongliang Liu
arXiv:2609.06074v2 Announce Type: replace
Abstract: Sparse optical flow provides stable inter-frame correspondence, playing a key role in Visual Odometry (VO) and Visual-Inertial Odometry (VIO). Clas...
By Yicheng Lin, Zhipeng Fei, Yuxiu Xu, WenDong Chen, Cong Li, Bin Han
arXiv:2511. 20332v3 Announce Type: replace-cross Abstract: This article trained a network for perceiving three-dimensional motion information of binocular vision target, which can provide real-time three-dimensional coordinate, velocity, and acceleration, and has a basic spatiotemporal perception capability.
By Jiazhao Shi, Pan Pan, Haotian Shi
arXiv:2608. 00508v1 Announce Type: cross Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research.
By Kai Geissler, Laurens M\"uller-Groh, Hans Meine
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv:2403. 06025v4 Announce Type: replace-cross Abstract: We introduce a new approach using computer vision to predict the land surface displacement from subsurface geometry images for Carbon Capture and Sequestration (CCS).
By Wei Chen, Yunan Li, Yuan Tian
arXiv:2607. 22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols.
By Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo\'nski, Tomasz Figatowski, Natalia Zieli\'nska
arXiv:2504. 06881v2 Announce Type: replace-cross Abstract: Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems, yet their reliance on multiplication-intensive computations poses challenges for deployment on resource-constrained devices.
By Mingbo Li, Liying Liu, Charles Wiranto, Ye Luo
arXiv:2507.13628v3 Announce Type: replace
Abstract: Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understand...
By Masahiro Ogawa, Qi An, Atsushi Yamashita
arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2503. 04500v3 Announce Type: replace-cross Abstract: Video understanding has largely relied on deep spatiotemporal architectures, including 3D convolutional networks and optical flow (OF) based models.
By Yu-Hsi Chen, Ching-Kai Lin, PingKong Huang, Chin-Tien Wu
TAPe+ML v3 is a compact computer vision system that uses a structured representation called TAPe to encode relationships among perceptual elements before recognition. The system employs a shared TAPe representation and a modular recognition architecture for tasks such as image classification, object detection, and instance segmentation, achieving strong performance with fewer than 100,000 parameters. Experiments show high mAP scores on COCO, strong classification accuracy on Imagenette and ImageNet-Real, and demonstrate benefits in video scene detection and distribution‑shift adaptation in an industrial pilot.
By Sergey Kurinov (Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia), Alexey Upatov (Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia)