arXiv AI By Musa Tur Farazi, K G Subarno Bithi

Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classification from Bangladeshi Radiographs

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computer Vision
Sep 25

Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

The paper introduces a dual‑input, multi‑task learning framework that jointly segments and classifies bone tumors by applying bidirectional cross‑modal attention between a lesion crop and the full radiograph. Using a YOLO‑based detector and a dual‑stream DenseNet121 architecture, the model fuses fine‑grained lesion detail with global anatomical context through a novel cross‑modal attention fusion strategy and hierarchical multi‑scale feature fusion. On the multi‑institutional Bone Tumor X‑ray Radiograph Dataset, the approach outperforms single‑input baselines, achieving a Dice coefficient of 0.896 and a macro‑averaged F1‑score of 0.928, with an AUC of 0.999 for malignant osteosarcoma.

By S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul
arXiv Machine Learning
Sep 21

Detection is solved, delineation is not: what governs tooth segmentation on panoramic radiographs

The study evaluates automatic tooth segmentation on panoramic radiographs using a large annotated corpus of 1,422 images and 42,142 tooth polygons. It finds that increasing input resolution improves boundary precision (mask mAP50‑95 rises from 0.656 to 0.717) while detection performance remains unchanged, and that architectural changes have minimal impact on in‑domain accuracy. Targeted interventions such as LoRA adaptation, promptable foundation models, and anatomical label assignment provide negligible gains, indicating that resolution and acquisition diversity should be prioritized over model novelty.

By Muhammad Rehan, Moaz Amjad, Syed Danial Ahmed, Mariam Adnan, Haider Ali