arXiv Computer Vision By Dania Khan, Nuzhat Aisha Shaikh, Asfina Hassan Juicy, Raiyun Kabir, S M Shahida, Taufiq Hasan

A Dual Cross-Attention Framework for Colposcopic CIN Grading and Swede Score Prediction Using a New Multi-Center Dataset

Read the original on arXiv Computer Vision →

The paper introduces a dual cross‑attention deep learning framework for automated grading of Cervical Intraepithelial Neoplasia (CIN) and prediction of Swede scores, using a newly released BUET Multi‑Center Colposcopy Dataset. The architecture fuses paired multimodal cervigrams and employs a custom composite loss to handle class imbalance, achieving 71.85% accuracy and 86.23% AUC‑ROC for three‑class CIN grading, and AUC‑ROC values between 75.7% and 88.4% for individual Swede score components. The total predicted Swede Score has a mean absolute error of 1.489, indicating potential for AI‑assisted colposcopy screening in resource‑limited settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Aug 26

RACR-MIL: Rank-aware contextual reasoning for weakly supervised grading of squamous cell carcinoma using whole slide images

RACR-MIL is a weakly‑supervised method for grading squamous cell carcinoma (SCC) from whole‑slide images, using an attention‑based multiple‑instance learning framework. It introduces a hybrid WSI graph to capture local tissue context and non‑local phenotypic dependencies, and applies rank‑ordering constraints on attention to prioritize higher‑grade tumor regions, mirroring pathologists’ diagnostic reasoning. The approach achieves state‑of‑the‑art performance, improving SCC grading accuracy by 3–9% over existing methods and up to 10% in tumor localization, and a pilot study showed pathologists reported increased grading efficiency in 60% of cases.

By Anirudh Choudhary, Mosbah Aouad, Krishnakant Saboo, Angelina Hwang, Jacob Kechter, Blake Bordeaux, Puneet Bhullar, David DiCaudo, Steven Nelson, Nneka Comfere, Emma Johnson, Olayemi Sokumbi, Jason Sluzevich, Leah Swanson, Dennis Murphree, Aaron Mangold, Ravishankar Iyer
arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv AI
Sep 7

Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs

The paper introduces the Cross‑Modal Triage Network (CMTN), a multimodal deep‑learning model that fuses a Swin Transformer V2 visual encoder with a PubMedBERT text encoder to perform severity‑based triage, pathology detection, and generate visual explanations for chest radiographs. Trained on 34,639 image‑text pairs from MIMIC‑CXR‑JPG, the CMTN achieves high ordinal agreement with reference labels (QWK = 0.9341) and excellent pathology detection (macro‑AUROC = 0.9970) while operating with 34 ms latency. However, a blinded clinical audit revealed low agreement with expert radiologists (QWK = 0.1399) and only modest spatial‑semantic concordance in heatmaps, underscoring the gap between algorithmic performance and clinical judgment.

By Zinah Ghulam, Richa Mittal, Eranga Ukwatta
arXiv Computer Vision
4d ago

Deep Learning-based Intelligent Diagnosis of Congenital Uterine Anomalies in 3D Ultrasound

arXiv:2609.15225v1 Announce Type: new Abstract: Objective: To develop an intelligent framework, termed CUA-Net, for the automated classification of congenital uterine anomalies (CUA) without requirin...

By Yueyue Xu, Yuhao Huang, Jiaxiao Deng, Yuanji Zhang, Haoming Zhang, Jiajia Qu, Shiying Zheng, Xiaomei Tang, Haining Chen, Chengcai Chen, Yiyi Wu, Xin Yang, Dong Ni
arXiv AI
Jun 30

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.

By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf