arXiv Computer Vision

Introduction to Computer Vision

arXiv AI
Jul 20

3D Motion Perception of Binocular Vision Target with PID-CNN

arXiv:2511. 20332v3 Announce Type: replace-cross Abstract: This article trained a network for perceiving three-dimensional motion information of binocular vision target, which can provide real-time three-dimensional coordinate, velocity, and acceleration, and has a basic spatiotemporal perception capability.

By Jiazhao Shi, Pan Pan, Haotian Shi
arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv Machine Learning
Jul 27

CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography

arXiv:2607. 22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols.

By Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo\'nski, Tomasz Figatowski, Natalia Zieli\'nska
arXiv Computer Vision
Sep 21

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

TAPe+ML v3 is a compact computer vision system that uses a structured representation called TAPe to encode relationships among perceptual elements before recognition. The system employs a shared TAPe representation and a modular recognition architecture for tasks such as image classification, object detection, and instance segmentation, achieving strong performance with fewer than 100,000 parameters. Experiments show high mAP scores on COCO, strong classification accuracy on Imagenette and ImageNet-Real, and demonstrate benefits in video scene detection and distribution‑shift adaptation in an industrial pilot.

By Sergey Kurinov (Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia), Alexey Upatov (Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia)