Self-supervised learning (SSL) can reduce the need for labelled medical images, but the choice of pretext objective remains unclear for lung ultrasound (LUS). Contrastive learning, masked reconstructi...
arXiv:2609.16551v1 Announce Type: new
Abstract: Self-supervised learning (SSL) can reduce the need for labelled medical images, but the choice of pretext objective remains unclear for lung ultrasound...
By Moein Heidari, Junbo Rao, Jai Choraria, Wenjin Chen, David J. Foran, Ilker Hacihaliloglu
arXiv:2609.09757v1 Announce Type: cross
Abstract: Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, gra...
By Hong Nguyen, Sean Foley, Christina Hagedorn, Yijing Lu, Sudarsana Reddy Kadiri, Dani Byrd, Shrikanth Narayanan
arXiv:2607. 29182v1 Announce Type: cross Abstract: Transcranial focused ultrasound (tFUS) is a non-invasive technique that delivers focused acoustic energy through the skull for neuromodulation and therapeutic applications.
By Minju Seol, Minjee Seo, Seonaeng Cho, Kyungho Yoon
Echo-E$^3$Net is an anatomy‑guided spatio‑temporal neural network designed to estimate left ventricular ejection fraction (LVEF) from ultrasound images. It uses a dual‑phase Endocardial Border Detector to locate end‑diastole and end‑systole landmarks and an Endocardial Feature Aggregator to fuse these landmarks with global deep‑feature descriptors for EF regression. The model achieves competitive accuracy on EchoNet‑Dynamic and EchoNet‑Pediatric datasets while using only 1.55 M parameters and 8.05 GFLOPs, enabling real‑time deployment on limited‑resource devices.
By Moein Heidari, Afshin Bozorgpour, AmirHossein Zarif-Fakharnia, Wenjin Chen, Dorit Merhof, David J. Foran, Jasmine Grewal, Ilker Hacihaliloglu
US-JEPA introduces a self‑supervised framework for ultrasound imaging that predicts masked latent representations instead of raw pixels, using a frozen, domain‑specific teacher to provide stable targets. This approach avoids the hyperparameter sensitivity and computational cost of traditional online teachers, enabling the student model to build upon the teacher’s semantic priors. The authors benchmark US‑JEPA against all publicly available ultrasound foundation models on UltraBench, showing competitive or superior performance across multiple organs and pathological conditions under linear probing.
By Ashwath Radhachandran, Vedrana Ivezi\'c, Shreeram Athreya, Corey W. Arnold, William Speier
arXiv:2511. 21325v2 Announce Type: replace-cross Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs.
By Ido Nitzan Hidekel, Gal lifshitz, Khen Cohen, Dan Raviv
FreqDINO++ is a frequency‑guided multi‑task routing vision foundation model designed for universal ultrasound analysis. It introduces a Multi‑task Routing Adapter for efficient task‑common and task‑specific integration, a Frequency‑aware Feature Enhancer to capture multi‑scale frequency characteristics, and a Task‑aligned Collaborative Decoder that promotes collaboration between dense and global prediction tasks. Experiments on large‑scale multi‑task and external single‑task ultrasound benchmarks show that FreqDINO++ outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, with promising generalization to unseen data.
By Qing Xu, Yixuan Zhang, Yue Li, Xiangjian He, Qian Zhang, Mainul Haque, Rong Qu, Wenting Duan, Jieyun Bai, Zhen Chen
arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.
By Fengrui Liu, Ruiyang Huang, Qijian Zheng, Yuanfang Wang, Feng Liu
arXiv:2607. 20136v1 Announce Type: cross Abstract: Slice-to-volume reconstruction (SVR) is the standard method for obtaining high-resolution (HR) 3D fetal brain volumes from motion-corrupted 2D MRI slice stacks acquired in multiple orientations.
By Busra Bulut, Maik Dannecker, Thomas Sanchez, Sara Neves Silva, Steven Jia, Jean-Baptiste Ledoux, Leo Pomar, Joanna Sichitiu, Yvan Gomez, Meriam Koob, Vincent Dunet, Maria Deprez, Guillaume Auzias, Francois Rousseau, Jana Hutter, Daniel Rueckert, Meritxell Bach Cuadra
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.
By Ludovic K. Tuncay (IRIT-SAMoVA), Etienne Labb\'e (IRIT-SAMoVA), Thomas Pellegrini (IRIT-SAMoVA)
arXiv:2606. 11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women receive no skilled sonography.
By Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus