arXiv:2606. 08206v1 Announce Type: cross Abstract: We present SegmentAnyTreeV2, a sensor- and platform-agnostic framework for semantic and instance segmentation of forest point clouds.
By Maciej Wielgosz, Stefano Puliti, Rasmus Astrup
arXiv:2609.26549v1 Announce Type: new
Abstract: Individual tree crown segmentation from aerial imagery underpins tree-level carbon accounting, biodiversity, and restoration monitoring at landscape sc...
By Thomas Pitts, Kunqi Li, Bin Liang
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision.
arXiv:2510. 09458v2 Announce Type: replace-cross Abstract: Interest in forestry automation is growing alongside rapid advances in deep learning.
By David-Alexandre Duclos, William Guimont-Martin, Gabriel Jeanson, Arthur Larochelle-Tremblay, Martine Lapointe, Th\'eo Defosse, Fr\'ed\'eric Moore, Philippe Nolet, Fran\c{c}ois Pomerleau, Philippe Gigu\`ere
The paper introduces BalSAM, a model that combines the Segment Anything Model (SAM) with Digital Surface Model (DSM) elevation data to improve tree crown instance segmentation from high‑resolution drone imagery. Experiments across boreal plantations, temperate forests, and tropical forests show that while off‑the‑shelf SAM does not beat a custom Mask R-CNN, fine‑tuning SAM end‑to‑end and incorporating DSM information yield promising results, especially for plantation sites.
By M\'elisande Teng, Arthur Ouaknine, Etienne Lalibert\'e, Yoshua Bengio, David Rolnick, Hugo Larochelle
arXiv:2609.24787v1 Announce Type: new
Abstract: Forest inventories increasingly rely on artificial intelligence (AI) models to derive forest attributes from large-scale 3D point clouds. Current model...
By Yuanwen Yue, Stefano Puliti, Damien Robert, Atakan Topalo\u{g}lu, Binbin Xiang, Maciej Wielgosz, Jan Dirk Wegner, Rasmus Astrup, Christian Rupprecht, Konrad Schindler
SelectAnyTree is a promptable instance segmentation model designed for 3D forest LiDAR point clouds, enabling users to delineate individual trees with a few clicks. The architecture comprises a sparse voxel scene encoder, a click‑to‑query prompt encoder, and a state‑space query decoder that produces tree masks in linear time, requiring only 19.4 M parameters. Across seven forest regions and an independent dataset, the model achieves a 79.9 % IoU for a single‑click target tree, outperforming existing promptable baselines and requiring the fewest clicks to reach accuracy targets.
By Trung Thanh Nguyen, Daniel Lusk, Kilian Gerberding, Janusch Vajna-Jehle, Tuan-Anh Vu, Duc Viet Le, Tu Vo, Phi Le Nguyen, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide, Julian Frey, Teja Kattenborn
arXiv:2608. 15790v1 Announce Type: new Abstract: Crevasse mapping from uncrewed aerial vehicle (UAV) imagery matters for glaciological research and for field safety in glaciated terrain.
By Steven Wallace, William D Harcourt, Richard Hann, Aiden Durrant, Somayajulu Sripada, Georgios Leontidis
arXiv:2608.28216v1 Announce Type: new
Abstract: Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is ab...
By Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Md. Sadman Haque, Mohd Ariful Haque
arXiv:2609.38714v1 Announce Type: new
Abstract: We describe our winning entry to the Waymo Open Dataset 2D Video Panoptic Segmentation Challenge. The task asks for a semantic class at every pixel of...
By Jinghan Yang
SenseFuse introduces a label‑free fusion approach that balances 2D image and 3D shape encoders for open‑vocabulary 3D instance segmentation. By selecting a scene‑level fusion weight through an adaptive, sensitivity‑based mechanism, it improves mask labeling accuracy across multiple datasets, recovering up to 93% of the potential gain from an oracle weight. The method demonstrates that image and shape encoders have complementary failure patterns, leading to higher instance AP in most evaluated settings.
By Euiseok Han, Tri Ton, Hwanhee Kim, Seungyeon Ryu, Chang D. Yoo
Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella