SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
arXiv:2510. 09458v2 Announce Type: replace-cross Abstract: Interest in forestry automation is growing alongside rapid advances in deep learning.
arXiv:2510. 09458v2 Announce Type: replace-cross Abstract: Interest in forestry automation is growing alongside rapid advances in deep learning.
arXiv:2608.31052v1 Announce Type: cross Abstract: Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or bui...
The paper introduces BalSAM, a model that combines the Segment Anything Model (SAM) with Digital Surface Model (DSM) elevation data to improve tree crown instance segmentation from high‑resolution drone imagery. Experiments across boreal plantations, temperate forests, and tropical forests show that while off‑the‑shelf SAM does not beat a custom Mask R-CNN, fine‑tuning SAM end‑to‑end and incorporating DSM information yield promising results, especially for plantation sites.
arXiv:2609.16773v1 Announce Type: new Abstract: Image segmentation remains challenging due to occlusions, poor lighting, and irregular structures. Although transformer-based methods achieve high accu...
The paper presents a deep learning approach for detecting woody clearing using bitemporal Sentinel‑2 imagery from New South Wales, Australia. By introducing a loss‑scaling coefficient, the authors align the model’s objective with end‑user metrics, boosting precision and recall. They further demonstrate zero‑shot transfer to woody regrowth and segmentation tasks, achieving significant error reductions and high F1 scores through image augmentation and generation techniques.
arXiv:2505. 04397v2 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored.
Image segmentation remains challenging due to occlusions, poor lighting, and irregular structures. Although transformer-based methods achieve high accuracy, they rely heavily on long-range spatial fea...
arXiv:2509.22404v2 Announce Type: replace Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
The paper presents a generative framework that estimates category-level 6D pose and 3D size of objects from a single RGB image, using score-based diffusion models to produce a multi-hypothesis pose distribution. It replaces costly likelihood pruning with a Mean Shift approach to isolate the mode as the final pose estimate, achieving state-of-the-art results on the REAL275 benchmark. The method also decouples detection from pose estimation, enabling robust zero-shot generalisation on the Wild6D dataset and extending naturally to video sequences by propagating the pose distribution over time.
arXiv:2607. 23371v1 Announce Type: cross Abstract: Vascular segmentation is a standard procedure for clinical diagnosis, yet the specific visual features determining model decisions remain poorly understood.
arXiv:2602. 10045v2 Announce Type: replace-cross Abstract: Current instance segmentation models achieve high performance on average predictions, but lack principled uncertainty quantification: their outputs are not calibrated, and there is no guarantee that a predicted mask is close to the ground truth.
Vascular segmentation is a standard procedure for clinical diagnosis, yet the specific visual features determining model decisions remain poorly understood. This paper investigates the visual cues Convolutional Neural Networks (CNNs) use to segment blood vessels across two distinct imaging domains: fluorescence microscopy and retinal fundus photography.