arXiv:2603. 13994v2 Announce Type: replace-cross Abstract: Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties.
By Hossein Adeli, Seoyoung Ahn, Andrew Luo, Mengmi Zhang, Nikolaus Kriegeskorte, Gregory Zelinsky
arXiv:2602. 24181v2 Announce Type: replace-cross Abstract: Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks.
By Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson, Ye Xia, Skanda Koppula, Andre Araujo, Joao Carreira, Niloy J. Mitra
DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D observations. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch fusion, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.
By Farkhat Almukhamedov, Sami Azirar, Hermann Blum
DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D data. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch techniques, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.
arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.
By Gabriel Jeanson, David-Alexandre Duclos, William Larriv\'ee-Hardy, No\'e Cochet, Mat\v{e}j Boxan, Anthony Desch\^enes, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity.
arXiv:2608.20720v1 Announce Type: new
Abstract: Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assum...
By Junqi Wu, Kaihua Tang, Xuanwen Chen, Hongzhi Li, Jianqiang Huang, Xian-Sheng Hua
arXiv:2606. 25128v1 Announce Type: cross Abstract: Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs.
By \"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understandin...
arXiv:2607. 21371v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories.
By Sung-Hoon Yoon, Hoyong Kwon, Changgyoon Oh, Kuk-Jin Yoon
arXiv:2609.38079v1 Announce Type: new
Abstract: Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and obje...
By Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DINOv3 provides strong structured visual representations, its lack of native textual alignment hinders its direct application to OVSS.