arXiv:2608.23024v1 Announce Type: new
Abstract: Counterfactual medical image generation aims to modify an existing image to reflect a hypothetical scenario in which certain characteristics of the ima...
By Andrea Posada, Wenke Karbole, Bach Ngoc Doan, Alexander Weers, Solmaz Abdolrahimzadeh, Maria Patsiamanidi, Kahkashan Haider, Vaishali Khare, Daniel Rueckert, Andrew Lotery, Sobha Sivaprasad, Martin J. Menten
arXiv:2609.12834v1 Announce Type: new
Abstract: Modelling how a disease progresses over time requires longitudinal imaging cohorts, which are scarce and small, whereas cross-sectional data -- one ima...
By Ifeoma Veronica Nwabufo, Julius Gervelmeyer, Sarah M\"uller, Philipp Berens
arXiv:2606. 04881v1 Announce Type: cross Abstract: Face aging plays an important role in long-term biometric analysis, cross-age identity verification, and forensic identity analysis.
By Yueying Zou, Peipei Li, Qianrui Teng, Dianyan Xu, Zekun Li
Co-Annotator is a clinical AI system that distills expert gaze and dictation into two guidance components: a gaze‑aligned Vision Transformer that highlights fixation‑aligned areas of interest (AOIs) and an ontology‑bounded vision‑language model that pre‑fills editable biomarker summaries for retinal OCT. In controlled studies, each modality independently improved diagnostic accuracy and biomarker generation, and when combined across two academic institutions, the system increased correct diagnoses per minute by 40% and reduced comment editing time by 67% without compromising accuracy.
By Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor
arXiv:2606. 15129v1 Announce Type: cross Abstract: Color fundus photography (CFP) is the mainstay for large-scale retinal screening, yet its diagnostic capacity is constrained by the lack of depth-resolved structural information.
By Zhuo Deng, Ruiheng Zhang, Ziheng Zhang, Weihao Gao, Yitong Li, Qian Wang, Lei Shao, Jiaoyue Dong, Zhixi Zeng, Lijian Fang, Haibo Wang, Xiaobin Lin, Tao Liu, Zhicheng Du, Zhengwei Zhang, Lin Yang, Zheng Gong, Xinyu Zhao, Zhenquan Wu, Fang Li, Zhiguang Zhou, Guoming Zhang, Sun Jing, Han Lv, Wenbin We, Lan Ma
arXiv:2608. 05938v1 Announce Type: cross Abstract: Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks under the assumption that this renders them anonymous.
By Attila Simk\'o
arXiv:2607. 21068v1 Announce Type: new Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reducing dependency on specialist-only assessment, for which neural network-based deep learning (DL) models have been widely utilized.
By Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee
The study introduces an explainable OCTA pipeline for Alzheimer’s disease phenotyping that combines vessel segmentation, layer‑specific biomarker extraction, and label‑free phenotyping with large language model (LLM) reporting. Using 117 images from 39 subjects, the segmentation models achieved high ROC‑AUC (0.916–0.970) and Dice scores (0.695–0.781), and six vascular biomarkers were used to create subject‑level profiles for exploratory clustering. LLMs (GPT, Gemini, Llama) produced measurement‑grounded reports evaluated for citation faithfulness and diagnostic caution, offering a transparent, non‑diagnostic link between retinal vascular data and Alzheimer’s research.
By Progga Paromita Dutta, Jeba Maliha, Md Rafiul Kabir
arXiv:2608.24723v1 Announce Type: new
Abstract: Retinal fundus photography is widely used for screening and monitoring ocular diseases, but many modern classification pipelines rely on deep latent re...
By Xiaoyan Li, Shixin Xu, Arvind Gupta, Huaxiong Huang
The paper presents an explainable diabetic retinopathy classification framework that leverages vision foundation models—DINOv2, CLIP, and Vision Transformer—combined with various transfer learning techniques such as full fine‑tuning, linear probing, and Low‑Rank Adaptation (LoRA). Using the ODIR dataset for internal validation and the APTOS dataset for external testing, DINOv2‑LoRA achieved the best internal AUROC (0.758) while DINOv2 and ViT full fine‑tuning reached the highest external AUROC (0.920). Explainability was assessed with Grad‑CAM and HiResCAM against expert‑annotated lesion masks from IDRiD, using Dice, IoU, and Pointing Game metrics, confirming that model attention aligns with clinically relevant retinal lesions.
By Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
arXiv:2603. 18846v3 Announce Type: replace-cross Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL).
By Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi, Philipp Berens
arXiv:2607. 05825v1 Announce Type: cross Abstract: Background.
By Fred Mutisya, Oscar Onyango, Sarah Sitati, Syokau Ilovi, Aeesha NJ Malik, Brenda W'mosi, Brian Makini, Jalemba Aluuvala, Josiah Onyango, Rachael Kanguha Mmene, Steven Wanyee