Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,329 stories · RSS feed

arXiv Computer Vision
1d ago

ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration

ActiveLang is an autonomous system that builds language‑annotated 3D maps by actively exploring environments. It uses a compact dual‑Gaussian representation to jointly reconstruct geometry, appearance, and open‑vocabulary semantics while adapting language features online. The planner selects informative viewpoints, reducing the number of observations and computational cost, and experiments on Replica and ScanNet++ show significant gains in 2D and 3D open‑vocabulary segmentation compared to existing baselines.

By Liyan Chen, Hairong Yin, Huangying Zhan, Yi Xu, Raymond A. Yeh, Philippos Mordohai
arXiv Computer Vision
1d ago

WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation

WAPR is a zero‑shot wide‑angle pose refinement model that can correct candidate 6D poses with rotational errors up to 90°, achieving fast inference (≤1 s per frame) and high throughput (≈25 detections per second). It leverages rotational symmetry priors to canonicalize pose targets and introduces the SA6D dataset, which augments 944 GSO scans into ~50 k object instances and ~2 M RGB‑D images. Experiments on seven BOP core datasets demonstrate that WAPR sets new state‑of‑the‑art performance for unseen‑object pose estimation in both fast and unconstrained settings.

By Yulin Wang, Mengting Hu, Hongli Li, Jianghao Zhou, Chen Luo
arXiv Machine Learning
1d ago

Lightweight and Versatile Learned Optimization by Recombination of Gradient History

The paper introduces a lightweight learned optimizer that recombines gradient history by averaging over disjoint time spans, reducing the prediction space to a single scalar coefficient per average shared across parameters. By progressively averaging older gradients, the method keeps memory usage low while maintaining independent contributions from long‑term history. A small 37k‑parameter network trained in under a GPU‑hour generalizes zero‑shot to unseen tasks, improving validation loss on BERT‑Tiny and GPT‑Tiny and boosting test accuracy over Adam on Vision Transformers and graph models with minimal FLOPs overhead.

By Minyoung Choi, Dalta Imam Maulana, Wanyeong Jung
arXiv Machine Learning
1d ago

Estimating Model-Level Membership Inference Vulnerability Without Reference Models

The paper introduces a method to estimate a model’s vulnerability to the Likelihood Ratio Attack (LiRA) without training reference models, using only the target model’s train and test loss distributions. It shows that LiRA’s per‑sample signal can be decomposed into a variance‑ratio term and a residual mean‑shift term, and that different loss‑distribution shapes dictate which reference‑free proxy to use. Two proxies are presented: the LOSS attack TNR for heavy‑tailed losses and the LOSS attack AUC for symmetric losses, both achieving low RMSE in predicting LiRA TPR across multiple architectures and datasets.

By Euodia Dodd, Nata\v{s}a Kr\v{c}o, Igor Shilov, Matthew Wicker, Yves-Alexandre de Montjoye
arXiv Computer Vision
1d ago

LiG-DETR: Local-in-Global Reassembly in Latent Space for Aerial Object Detection

LiG-DETR introduces a Global-Local Reassembly framework for aerial object detection that captures high‑fidelity local features before compression and integrates them into a unified end‑to‑end DETR decoder. The method uses a shared encoder to extract both global and locally magnified features, reassembles the local features according to their spatial positions, and employs Context‑Preserved Selective Reassembly and Density‑Aware Adaptive Query Allocation to reduce redundant computation. Experiments demonstrate significant improvements on small and medium objects while maintaining strong performance on large objects, with favorable accuracy–efficiency trade‑offs and better cross‑domain generalization.

By Yupeng Zhang, Fangzhuo Gao, Juntao Cheng, Ziyi Zhao, Liang Wan, Ruize Han