arXiv:2606. 03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design.
By Wei Ding, Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Yu Wang
arXiv:2609.13233v1 Announce Type: new
Abstract: Thin structures such as tree branches are among the hardest cases for stereo matching: a branch is only a few pixels wide, the background is cluttered,...
By Yida Lin, Bing Xue, Mengjie Zhang, Sam Schofield, Richard Green
arXiv:2606. 02092v1 Announce Type: cross Abstract: Semantic segmentation of remote sensing imagery requires models that capture both global context and local detail under tight computational budgets.
By \"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel
arXiv:2608. 10989v1 Announce Type: cross Abstract: Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands.
By Hongsen Cao, Mona Jaber, Shanxin Yuan, Ahmed Sayed
The paper introduces Successive Capacity Growth (SCG), a method for adaptively expanding Vision Transformer encoders in Joint-Embedding Predictive Architectures (JEPAs). SCG starts with a minimal encoder and incrementally increases width or depth based on a task‑agnostic test‑and‑verify mechanism, while a Sketched Isotropic Gaussian Regularizer (SIGReg) keeps learned semantic dimensions independent. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves up to 20.3% better prediction loss than fixed small baselines and 23% better than fixed large models, with far greater parameter efficiency and no false‑positive expansions.
By Frederik Berenz
The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.
By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv:2609.26549v1 Announce Type: new
Abstract: Individual tree crown segmentation from aerial imagery underpins tree-level carbon accounting, biodiversity, and restoration monitoring at landscape sc...
By Thomas Pitts, Kunqi Li, Bin Liang
arXiv:2607. 16012v1 Announce Type: cross Abstract: Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation.
By Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim
The paper introduces CP‑BG‑Bench, a paired‑view evaluation framework for Cell Painting vision encoders that fixes a central cell across four matched views (raw crop, segmented, and density‑augmented variants). Using this framework on three datasets and three encoders, the authors show that standard single‑metric rankings (e.g., replicate mAP) vary systematically across protocols, revealing disagreements along axes of cell versus background, morphology versus context, and within‑study versus across‑batch performance. The study demonstrates that segmented views can outperform crops in certain tasks and that background‑driven gains are largely determined by experimental design rather than encoder choice.
By Tim Treis, Nikita Moshkov, Johan Fredin Haslum, Shantanu Singh, Fabian J. Theis
Semantic segmentation of remote sensing imagery requires models that capture both global context and local detail under tight computational budgets. Prior work typically optimizes for one of these axes: attention for global context, convolution for local detail, or compactness for efficiency.
arXiv:2503.10685v3 Announce Type: replace
Abstract: Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited d...
By Brun\'o B. Englert, Gijs Dubbelman
LiteViLNet is a lightweight RGB‑geometry fusion network for road segmentation that uses a MobileNetV3 RGB encoder and a tiny depth‑wise‑separable geometry encoder. Its multi‑scale fusion module enhances modality‑specific features, performs cross‑modal interaction, and applies adaptive gating, while a depth‑wise large‑kernel bridge expands contextual support with minimal overhead. The U‑Net‑style decoder is trained with deep supervision, achieving state‑of‑the‑art performance on KITTI and ORFD benchmarks and running at up to 68.73 FPS on a Jetson Orin NX with TensorRT FP16.
By Daojie Peng, Bingtao Wang, Fulong Ma, Liang Zhang, Jun Ma