arXiv Computer Vision By Gyeonggwan Lee, Eunsoo Im, Seunghwan Hong, Junghun Suh

SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence

Read the original on arXiv Computer Vision →

SCCM (Spherically Consistent Coarse Matching) improves dense feature correspondence on equirectangular projection (ERP) imagery by correcting topological, metric, and area distortions at the coarse matching stage. It augments a chart‑naive cross‑attention matcher with spherical positional attention and area‑aware covisibility priors, raising PCK@1° from 0.229 to 0.275 on Matterport3D while keeping the refiner unchanged. In the RoMa V1 framework, SCCM outperforms ERP‑native EDM and an ERP‑retrained RoMa V1, and it transfers zero‑shot to Stanford2D3D and outdoor Holo360D.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 20

MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors

Classical image correspondence is solved at the level of sparse keypoints or dense pixels, but the systems that consume these matches - object-level mapping, topological navigation, scene-graph maintenance - reason about whole objects. Recent work narrows this gap by matchng directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are obtained by pooling features from large 3D foundation models over the masks.

arXiv Computer Vision
Aug 27

TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection

TDFNet introduces a Tri-projection Deformable Fusion Network that uses equirectangular, cube map, and tangent projections to mitigate geometric distortions in panoramic salient object detection. It incorporates a cross-projection deformable attention module for geometry-aware sampling and a latitude-guided fusion module that balances ERP and CMP features using spherical latitude priors. The network’s three-branch encoding preserves global continuity, local detail, and boundary precision, improving detection performance over existing projection-based methods.

By Qiangqiang Zhou, Jiacong Yu, Jiawei Xu, Yong Chen, Xin Huang, Ping Li