Introduction to Computer Vision
arXiv:2609.39627v1 Announce Type: new Abstract: This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organi...
arXiv:2511. 20332v3 Announce Type: replace-cross Abstract: This article trained a network for perceiving three-dimensional motion information of binocular vision target, which can provide real-time three-dimensional coordinate, velocity, and acceleration, and has a basic spatiotemporal perception capability.
arXiv:2609.39627v1 Announce Type: new Abstract: This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organi...
The paper introduces Temporal Residual Neural Radiance Fields for reconstructing dynamic human bodies from monocular video. It builds a temporal residual field independent of MLPs, reduces trainable parameters, speeds up rendering, and employs a multi‑dimensional loss to improve pixel‑level accuracy. Experiments show higher PSNR and SSIM than recent methods while being roughly 780 times faster than Anim‑NeRF and Neural Body.
arXiv:2609.06074v2 Announce Type: replace Abstract: Sparse optical flow provides stable inter-frame correspondence, playing a key role in Visual Odometry (VO) and Visual-Inertial Odometry (VIO). Clas...
arXiv:2512. 18046v2 Announce Type: replace Abstract: Unmanned Aerial Vehicles, commonly known as, drones pose increasing risks in civilian and defense settings, demanding accurate and real-time drone detection systems.
arXiv:2603. 25157v3 Announce Type: replace-cross Abstract: Recent vision backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress on image recognition.
arXiv:2609.24403v1 Announce Type: new Abstract: Biological visual systems achieve continuous, low-latency motion perception by processing sparse, asynchronous spiking signals, enabling real-time trac...
arXiv:2607.05205v2 Announce Type: replace Abstract: Fast and reliable motion detection is essential for machine vision and autonomous systems operating in dynamic environments. This work integrates e...
Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.
TDFNet introduces a Tri-projection Deformable Fusion Network that uses equirectangular, cube map, and tangent projections to mitigate geometric distortions in panoramic salient object detection. It incorporates a cross-projection deformable attention module for geometry-aware sampling and a latitude-guided fusion module that balances ERP and CMP features using spherical latitude priors. The network’s three-branch encoding preserves global continuity, local detail, and boundary precision, improving detection performance over existing projection-based methods.
arXiv:2608. 13513v1 Announce Type: cross Abstract: Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers.
arXiv:2609.16637v1 Announce Type: cross Abstract: Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larg...
arXiv:2603. 25157v2 Announce Type: replace Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress, enabling unified modeling across images, text, and beyond.