The paper introduces TSR-ITNR, a two‑stage, self‑supervised framework for hyperspectral image super‑resolution that fuses high‑resolution multispectral and low‑resolution hyperspectral data. Stage 1 refines an implicit Tucker representation using a low‑rank spatial tensor and spectral basis, enhanced by a pretrained denoiser, to capture fine spatial details and spectral correlations. Stage 2 applies parameter‑free calibration to extract complementary corrections from both observations, preserving geometry and ensuring orthogonal complementarity, leading to superior reconstruction quality demonstrated on benchmark datasets and improved downstream segmentation performance.
By Liqian Yang, Xingchi Chen, Xinfeng Gui, Xiangyong Cao, Qianxin Yi
The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.
By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran
arXiv:2408. 15344v2 Announce Type: replace Abstract: Many scientific and engineering problems involve observing a common phenomenon through multiple heterogeneous sensors or measurement modalities.
By George A. Kevrekidis, Eleni D. Koronaki, Dimitris G. Giovanis, Yannis G. Kevrekidis
arXiv:2606. 31609v1 Announce Type: cross Abstract: Radar sensors provide reliable perception under adverse weather and lighting conditions, but their sparse, noisy, and weakly semantic measurements make dense semantic segmentation challenging.
By Ali Zia, Muhammad Umer Ramzan, Abdelwahed Khamis, Usman Ali, Abdul Rehman
MEOX is a compact multimodal masked autoencoder designed for Earth Observation that uses a 2.939 million‑parameter encoder and 3.115 million total parameters. It incorporates sensor‑specific adapters, explicit validity signals, and a shared sparse‑expert block to maintain modality‑dependent processing before a learned patch‑wise fusion, followed by fourteen encoder blocks that process a single spatial sequence with four metadata tokens. Pretrained on 1.228 million MMEarth64 samples, MEOX achieves strong performance on GEO‑Bench tasks, surpassing prior CSMoE results, and demonstrates effective sensor‑flexible representation learning with a modest parameter budget.
By Mohanad Albughdadi
arXiv:2606. 04280v1 Announce Type: cross Abstract: Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful latent geometry remain incompletely understood.
By Justinas Zaliaduonis, Patrick Putzky, Till Richter, Sergios Gatidis
arXiv:2607. 22705v1 Announce Type: cross Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations.
By Anuraag Gadehothur Karnam, Tarunesh Sathish
arXiv:2608. 15647v1 Announce Type: cross Abstract: Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult.
By Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin
arXiv:2607. 01902v1 Announce Type: cross Abstract: Reliable confidence estimates are essential in semantic segmentation, especially in safety-critical settings where overconfident errors can mislead downstream decisions.
By Tristan Kirscher (ICube), Kim-Celine Kahl (DKFZ), Balint Kovacs (DKFZ), Maximilian R. Rokuss (DKFZ), Klaus Maier-Hein (DKFZ), Xavier Coubez (ICube), Philippe Meyer (ICube), Sylvain Faisan (ICube)
arXiv:2608. 03319v1 Announce Type: cross Abstract: Future integrated sensing and communication (ISAC) architectures separate the sensing entity (SE) that acquires measurements from the sensing function (SF) that performs inference, creating a need for compact, task-oriented feedback on the SE-SF interface.
By Shiv Shankar, Radha Krishna Ganti, J Klutto Milleth
Reliable confidence estimates are essential in semantic segmentation, especially in safety-critical settings where overconfident errors can mislead downstream decisions. Yet modern segmentation models often remain miscalibrated.
The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and uses context-aware sampling to build compact, scene‑specific training sets for ship detection. By enriching frames with weather and solar metadata and applying diversity sampling, the method reduces 40,718 candidate frames to just 220 for annotation, achieving a 99.5% reduction. Fine‑tuned YOLOv26‑m on this curated subset attains high detection performance (AP50 ≈ 94.8% and AP50‑95 ≈ 75.1%).