arXiv Computer Vision

Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

The paper introduces a dual‑input, multi‑task learning framework that jointly segments and classifies bone tumors by applying bidirectional cross‑modal attention between a lesion crop and the full radiograph. Using a YOLO‑based detector and a dual‑stream DenseNet121 architecture, the model fuses fine‑grained lesion detail with global anatomical context through a novel cross‑modal attention fusion strategy and hierarchical multi‑scale feature fusion. On the multi‑institutional Bone Tumor X‑ray Radiograph Dataset, the approach outperforms single‑input baselines, achieving a Dice coefficient of 0.896 and a macro‑averaged F1‑score of 0.928, with an AUC of 0.999 for malignant osteosarcoma.

arXiv Machine Learning
Aug 17

XtraLight-MedMamba for Classification of Neoplastic Tubular Adenomas

arXiv:2602. 04819v5 Announce Type: replace-cross Abstract: Accurate risk stratification of precancerous polyps during routine colonoscopy screening is a key strategy to reduce the incidence of colorectal cancer (CRC).

By Aqsa Sultana, Rayan Afsar, Ahmed Rahu, Surendra P. Singh, Brian Shula, Brandon Combs, Derrick Forchetti, Vijayan K. Asari
arXiv Machine Learning
Jul 14

BiLoG-Net: A Bi-Context Location-Guided Network for Breast Mass Segmentation and Malignancy Classification in Mammography

arXiv:2607. 10188v1 Announce Type: cross Abstract: Breast cancer remains the most commonly diagnosed malignancy among women worldwide, yet accurate detection and characterization of breast masses in mammography remain challenging due to subtle intensity variations, heterogeneous tissue densities, and indistinct lesion boundaries that complicate radiological interpretation.

By Abu Fatema Mohammad Abdun Noor, Md Imam Ahasan, Md Samiul Ahasan, Kah Ong Michael Goh, S M Hasan Mahmud, Raihana Zannat
arXiv Machine Learning
Aug 26

RACR-MIL: Rank-aware contextual reasoning for weakly supervised grading of squamous cell carcinoma using whole slide images

RACR-MIL is a weakly‑supervised method for grading squamous cell carcinoma (SCC) from whole‑slide images, using an attention‑based multiple‑instance learning framework. It introduces a hybrid WSI graph to capture local tissue context and non‑local phenotypic dependencies, and applies rank‑ordering constraints on attention to prioritize higher‑grade tumor regions, mirroring pathologists’ diagnostic reasoning. The approach achieves state‑of‑the‑art performance, improving SCC grading accuracy by 3–9% over existing methods and up to 10% in tumor localization, and a pilot study showed pathologists reported increased grading efficiency in 60% of cases.

By Anirudh Choudhary, Mosbah Aouad, Krishnakant Saboo, Angelina Hwang, Jacob Kechter, Blake Bordeaux, Puneet Bhullar, David DiCaudo, Steven Nelson, Nneka Comfere, Emma Johnson, Olayemi Sokumbi, Jason Sluzevich, Leah Swanson, Dennis Murphree, Aaron Mangold, Ravishankar Iyer
arXiv Computer Vision
Sep 25

OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis

OncoVision is a privileged‑information training framework that learns from mammography images and clinical data during training but performs inference using only mammographic images. It employs an attention‑based encoder‑decoder to jointly segment masses, calcifications, axillary findings, and breast tissue, and predicts ten structured clinical features such as BI‑RADS. Two late‑fusion strategies (Independent and Dependent) integrate imaging, radiomic, and clinical information to improve diagnostic precision, and a retrospective multi‑reader study showed higher diagnostic confidence, reduced reading time, and segmentation accuracy comparable to or better than radiologists.

By Istiak Ahmed, Galib Ahmed, K. Shahriar Sanjid, Md. Tanzim Hossain, Md. Nishan Khan, Md. Misbah Khan, Md. Arifur Rahman, Sheikh Anisul Haque, Sharmin Akhtar Rupa, Mohammed Mejbahuddin Mia, Mahmud Hasan Mostofa Kamal, Md. Mostafa Kamal Sarker, M. Monir Uddin
arXiv AI
Jul 8

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

arXiv:2607. 06309v1 Announce Type: cross Abstract: Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities.

By Aysan Ghayouri Pirsoltan, Shima Babakordi, Mohammad Reza Mohammadi
arXiv AI
Aug 17

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

arXiv:2608. 13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5).

By Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li