arXiv AI

Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients

arXiv:2606. 16210v1 Announce Type: new Abstract: Learned representations in intelligent sensing systems are often evaluated by reconstruction fidelity or downstream prediction accuracy, but these criteria do not specify which latent distinctions are justified by the sensing process.

arXiv Computer Vision
Sep 7

Learning Spatial-Spectral Refinement and Calibrating Complementary Observations for Hyperspectral Image Super-Resolution

The paper introduces TSR-ITNR, a two‑stage, self‑supervised framework for hyperspectral image super‑resolution that fuses high‑resolution multispectral and low‑resolution hyperspectral data. Stage 1 refines an implicit Tucker representation using a low‑rank spatial tensor and spectral basis, enhanced by a pretrained denoiser, to capture fine spatial details and spectral correlations. Stage 2 applies parameter‑free calibration to extract complementary corrections from both observations, preserving geometry and ensuring orthogonal complementarity, leading to superior reconstruction quality demonstrated on benchmark datasets and improved downstream segmentation performance.

By Liqian Yang, Xingchi Chen, Xinfeng Gui, Xiangyong Cao, Qianxin Yi
arXiv Machine Learning
1d ago

Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions

The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.

By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran
arXiv Computer Vision
Sep 7

MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation

MEOX is a compact multimodal masked autoencoder designed for Earth Observation that uses a 2.939 million‑parameter encoder and 3.115 million total parameters. It incorporates sensor‑specific adapters, explicit validity signals, and a shared sparse‑expert block to maintain modality‑dependent processing before a learned patch‑wise fusion, followed by fourteen encoder blocks that process a single spatial sequence with four metadata tokens. Pretrained on 1.228 million MMEarth64 samples, MEOX achieves strong performance on GEO‑Bench tasks, surpassing prior CSMoE results, and demonstrates effective sensor‑flexible representation learning with a modest parameter budget.

By Mohanad Albughdadi
arXiv AI
Aug 18

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

arXiv:2608. 15647v1 Announce Type: cross Abstract: Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult.

By Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin
arXiv Machine Learning
Jul 3

Rethinking Post-Hoc Calibration in Semantic Segmentation

arXiv:2607. 01902v1 Announce Type: cross Abstract: Reliable confidence estimates are essential in semantic segmentation, especially in safety-critical settings where overconfident errors can mislead downstream decisions.

By Tristan Kirscher (ICube), Kim-Celine Kahl (DKFZ), Balint Kovacs (DKFZ), Maximilian R. Rokuss (DKFZ), Klaus Maier-Hein (DKFZ), Xavier Coubez (ICube), Philippe Meyer (ICube), Sylvain Faisan (ICube)
arXiv Machine Learning
Aug 5

Task-Oriented Candidate-Latent Feedback for Coarse-to-Fine Sensing in Distributed OFDM-ISAC Networks

arXiv:2608. 03319v1 Announce Type: cross Abstract: Future integrated sensing and communication (ISAC) architectures separate the sensing entity (SE) that acquires measurements from the sensing function (SF) that performs inference, creating a need for compact, task-oriented feedback on the SE-SF interface.

By Shiv Shankar, Radha Krishna Ganti, J Klutto Milleth
Hugging Face Trending Papers
Sep 24

Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

The paper introduces a pipeline that localizes frames from historical PTZ maritime video onto a reference panorama and uses context-aware sampling to build compact, scene‑specific training sets for ship detection. By enriching frames with weather and solar metadata and applying diversity sampling, the method reduces 40,718 candidate frames to just 220 for annotation, achieving a 99.5% reduction. Fine‑tuned YOLOv26‑m on this curated subset attains high detection performance (AP50 ≈ 94.8% and AP50‑95 ≈ 75.1%).