Robust Promptable Video Object Segmentation
arXiv:2605.12006v2 Announce Type: replace Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deploymen...
Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.
arXiv:2605.12006v2 Announce Type: replace Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deploymen...
arXiv:2605.19410v2 Announce Type: replace Abstract: Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-ho...
arXiv:2606.18623v2 Announce Type: replace Abstract: Gaussian segmentation is usually posed as transferring object knowledge from 2D foundation models into a 3D representation. This leaves a fundament...
arXiv:2606.20752v2 Announce Type: replace Abstract: Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recen...
HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.
arXiv:2609.34895v2 Announce Type: replace Abstract: Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this li...
arXiv:2606.21030v2 Announce Type: replace-cross Abstract: Diffusion-based image compression has achieved strong perceptual quality at ultra-low bitrates. However, existing codecs are often tied to sp...
arXiv:2608.24486v2 Announce Type: replace-cross Abstract: Reliable clot-volume quantification and subsequent risk assessment in pulmonary embolism depend on precise segmentation of emboli on computed...
arXiv:2609.37027v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-...
arXiv:2512.18991v3 Announce Type: replace-cross Abstract: Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds...
arXiv:2609.37298v1 Announce Type: new Abstract: Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classificat...
arXiv:2609.37631v1 Announce Type: new Abstract: Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work sh...
arXiv:2609.37659v1 Announce Type: cross Abstract: There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit....
arXiv:2609.36875v1 Announce Type: new Abstract: Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Any...
arXiv:2609.36891v1 Announce Type: new Abstract: Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approac...
arXiv:2609.36940v1 Announce Type: new Abstract: Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are esse...
arXiv:2606.08866v2 Announce Type: replace Abstract: CNN-based semantic segmentation networks usually rely on context heads such as ASPP, PPM, or attention modules to enlarge the receptive field. Thes...
The paper proposes Layer-Informed Fine-Tuning (LIFT), a method that identifies and updates only the most functionally critical layers of large language models (LLMs) using a bottleneck identification mechanism based on sensitivity analysis. By focusing on layers that handle conceptualization, reasoning, and textualization, LIFT aims to accelerate training and enhance performance on reasoning tasks. Experiments demonstrate that this selective fine-tuning approach both speeds up the training process and yields significant performance gains.
ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.
arXiv:2607.28623v2 Announce Type: replace-cross Abstract: We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for who...