Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook t...
Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model...
Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general o...
BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.
By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
The paper presents a self‑supervised approach to learning representations of auroral emission spectra using a 1D Vision Transformer trained with a masked autoencoder on 223,000 unlabelled spectra. The pretrained model recovers key emission‑line intensity ratios with high accuracy (R² = 0.91) and, with a single linear probe, matches expert‑designed feature classifiers. When fine‑tuned, it surpasses the previous supervised auroral classifier (macro‑AP 88.5 vs. 77.8) and achieves a 0.870 mAP, outperforming a model trained from scratch by 0.159 using only 10 % of the labels, while attribution reveals reliance on N₂⁺ bands.
By Matthieu Le Lain, Ga\"el Cessateur, S\'ebastien Lef\`evre
The paper introduces AURA, a meta‑learning framework that learns a low‑dimensional latent state‑space model for the evolution of optimal model parameters under distribution shift. Online adaptation is performed via extended Kalman filtering in this latent space, followed by reconstruction of full model parameters through a learned lifting map, enabling efficient single‑step updates. Experiments on neural wireless receivers and non‑stationary image classification show that AURA improves adaptation speed, accuracy, and computational efficiency compared to existing online learning and Bayesian filtering baselines.
By Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
The paper introduces the Phase-Coherent Transformer (PCT), a complex-valued architecture that replaces traditional softmax attention with a real-valued, smooth gate applied to L2-normalised query-key similarities. PCT eliminates token competition, preserving phase information across layers, and demonstrates strong generalisation on a variety of mid-scale benchmarks, outperforming both standard softmax Transformers and other complex-valued counterparts. Experiments confirm that the gate design is essential: preserving negatively aligned phase components is crucial for performance, while violating these conditions leads to degradation or collapse on long-range tasks.
By Leona Hioki
The paper evaluates whether incorporating the hierarchical structure of Tironian notes can improve automatic recognition of this complex Latin shorthand system. Experiments compare flat classifiers (ResNet18, ConvNeXt, Swin, ViT) with hierarchy‑aware models (HD‑CNN and routing approaches) on handwritten and manuscript samples, with and without few‑shot adaptation. Results show that hierarchical models outperform flat ones when no adaptation is applied, but flat models surpass them after few‑shot adaptation, indicating that hierarchy can aid recognition under non‑adapted conditions.
The paper reports a controlled study of self‑supervised learning (SSL) objectives for image and video pretraining under limited data, architecture, and compute budgets. It compares contrastive, reconstruction, feature‑prediction, and diffusion methods, finding that DINOv2‑style pretraining delivers the best overall performance. Combining DINOv2 with video SSL objectives such as VideoMAE improves image classification and segmentation but harms video tracking and camera‑pose estimation, highlighting a trade‑off between semantic and geometric learning.
By Brun\'o B. Englert, Gijs Dubbelman
The paper introduces SignShift, a framework for visual-only sentence-level segmentation of continuous sign language videos. It uses a Temporal Difference Module that captures frame-to-frame feature variations across full-frame, facial, and hand cues, and a Segment Count Prediction module to guide boundary selection. Experiments on benchmark datasets show that SignShift outperforms existing methods, demonstrating its effectiveness for this challenging task.
By Bowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu, Kuizhuang Liu, Zhiwei Jiang, Lei Xie
MDSkin-Net is a multi‑task skin lesion analysis framework that integrates Pattern Analysis priors into a hybrid CNN‑Transformer architecture. It introduces a Pattern Analysis‑Guided Attention Module (PAGAM) with improved Efficient Channel Attention, Multi‑Scale Spatial Attention, and Biased Asymmetry Attention, along with a multi‑scale spatial alignment regularization that uses segmentation masks as soft supervision. Trained only on the ISIC 2017 training split, the model achieves high segmentation and classification performance on multiple datasets, demonstrating strong zero‑shot generalization across different cohorts.
By Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas, Nikolaos Papanikolopoulos
The paper introduces a penalized distributionally robust optimization framework that allows an adversary to choose any distribution while incurring a Wasserstein penalty for deviating from the empirical distribution. It shows that the adversary’s problem can be reformulated as optimizing transport maps that push empirical samples to adversarial ones, proving that optimal maps are cyclically monotone. The authors argue that standard per-sample adversarial training violates this property and propose two remedies—multi-start particle ascent and input-convex neural network parameterization—to enforce cyclical monotonicity, demonstrating improved robustness and generalization in experiments on regression, image classification, and control tasks.
By Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse \c{S}en, Marco Cuturi, Daniel Kuhn
The paper introduces MR‑DiffuSR, a 3‑D latent diffusion framework that uses high‑resolution T1w structural priors to guide super‑resolution of thick‑slice FLAIR MRI scans. By applying cross‑modality structural swin attention and a mixed‑scale degradation strategy, the method avoids hallucinations and remains robust across varying slice thicknesses. On ADNI datasets, MR‑DiffuSR outperforms CNN and 2‑D diffusion baselines, achieving high PSNR, SSIM, and low LPIPS, and maintains strong white‑matter hyperintensity segmentation performance even at 7 mm equivalent slice thickness.
By Haoyu Lan, Jiazhen Zhang, John Onofrey, Bino Varghese, Nasim Sheikh-Bahaei, Arthur W. Toga, Jeiran Choupan
The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.
By Amir Zamani, Zeinab Ghasemi-Naraghi
The paper introduces PIMDE, a self‑supervised monocular depth estimation framework that decomposes input images into perceptual feature maps, each encoding a specific visual cue. Separate depth branches process these maps to produce individual depth estimates, which are then fused explicitly. Experiments on the KITTI benchmark show that PIMDE matches the accuracy of existing self‑supervised methods while offering clearer insight into how each perceptual cue contributes to depth prediction.
By Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis
arXiv:2609.31456v1 Announce Type: new
Abstract: Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hyp...
By Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers, Srinivasan Parthasarathy
arXiv:2609.31461v1 Announce Type: new
Abstract: Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a k...
By Xinxin Wang, Liam Hazan, Jing Li, Simona Rabinovici-Cohen, Xiaojuan Li, Mingrui Yang
arXiv:2609.31573v1 Announce Type: new
Abstract: Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and c...
By Ziyao Shang, Pouya Sadeghi, Letian Jiang, Alexander Wong, Sirisha Rambhatla
arXiv:2609.31452v1 Announce Type: cross
Abstract: Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and...
By Domen Tabernik, Peter Nimac, Jan Jeri\'cevi\'c, Danijel Sko\v{c}aj, Andrej Gams
arXiv:2406.09896v3 Announce Type: replace
Abstract: Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safe...
By Brun\'o B. Englert, Fabrizio J. Piva, Tommie Kerssies, Daan de Geus, Gijs Dubbelman