arXiv AI

Potential of Artificial Intelligence Algorithms for Identification of Relevant Diagnostic and Prognostic Biomarkers of Early-Stage Liver Cancer

This study investigates deep learning and explainable AI methods to diagnose hepatocellular carcinoma (HCC) and identify diagnostic and prognostic biomarkers across five disease stages using a transcriptomic dataset built via semi‑supervised learning. The best model used 15 genes selected by SelectKBest, achieving 90.74% accuracy, while a 20‑gene model had the lowest loss of 0.3187. SHAP‑based XAI highlighted DNAJB14 as the most influential gene, and functional validation showed that inhibiting DNAJB14 reverses key malignant traits of HCC cells.

arXiv AI
Sep 16

Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma Prediction

The article introduces a new dataset for hepatocellular carcinoma (HCC) prediction, comprising 770 patient samples with 11,150 gene expression levels each, categorized into five classes from normal tissue to various HCC stages. The dataset was constructed using XGBoost and semi‑supervised learning on three existing genomic biomarker datasets, leveraging their labels to generate new annotations. The resulting XGBoost model achieved a 96.5% classification accuracy during the semi‑supervised training process.

By Ahmed Ammar Kubba, Manar Abu Talib, Jibran Sualeh Muhammad, Ali Bou Nassif, Abdalla Sayed Mohamed, Darko Castven, Jens U. Marquardt
arXiv Machine Learning
Aug 18

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

The paper introduces AMPLIFAI, the first public dataset of multiphase abdominal CT scans annotated with LI-RADS categories and segmented for three key LI-RADS features: arterial phase hyperenhancement, washout, and enhancing capsule. It outlines the dataset’s composition, curation process, and annotation pipeline following the Datasheets for Datasets format to promote transparency and reproducibility. The dataset aims to support the development of AI models for automated hepatocellular carcinoma diagnosis using the biopsy‑free, imaging‑based LI‑RADs framework.

By Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi, Jana Delfino, James Tonascia, Jade Wong-You-Cheong, Barton Lane, Joseph Chirico, Jeffrey D. Hirsch, Ang Li, Heng Huang, Florence X. Doo
arXiv AI
Jul 7

Semantic Segmentation-Driven Image-Level Diagnosis of Liver Cancers in Hematoxylin and Eosin Histopathology Images

arXiv:2607. 03253v1 Announce Type: cross Abstract: As hematoxylin & eosin (H&E) staining constitutes the primary entry point in routine diagnostic workflows, computer-aided diagnosis from whole-slide H&E images is of particular clinical relevance.

By Ivica Kopriva, Dario Sitnik, Arijana Pacic, Karolina Krstanac, Irena Veliki Dalic, Marijana Popovic Hadzija
arXiv AI
Jul 7

CaresAI at SMM4H-HeaRD 2026: Predicting TNM Staging

arXiv:2607. 03466v1 Announce Type: cross Abstract: This study aims to predict Tumor, Node, and Metastasis (TNM) stage labels independently, with the Cancer Genome Atlas (TCGA) pathology report as the sixth shared task of SMM4H-HeaRD 2026.

By Joseph Itopa Abubakar, Jorge Jarme, Favour Igwezeke, Mary Adewunmi
arXiv Machine Learning
Jun 2

Early Prediction of Liver Cirrhosis Up to Two Years in Advance: A Machine Learning Study Benchmarking Against the FIB-4 and APRI Scores

arXiv:2601. 00175v2 Announce Type: replace Abstract: Objective: Develop and evaluate machine learning (ML) models for predicting incident liver cirrhosis (LC) one and two years prior to diagnosis using routinely collected electronic health record (EHR) data and benchmark their performance against the FIB-4 and APRI clinical scores.

By Zhuqi Miao, Ahmed G Qasem, Sujan Ravi, Jason T. Cheng, Abdulaziz Ahmed, Courtney W. Houchen, Sumayah Abed, Dilorom Azimdjanovna Zuparova, Abdulaziz Ahmed
arXiv AI
Sep 10

The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]

The study evaluates the use of default decision thresholds (t=0.50) in multi‑label enzyme commission (EC) number prediction across 14,096 compounds and six EC classes. It finds a high mean accuracy of 77.16% but low macro F1 (0.3976) and macro recall (0.3872), indicating severe class‑imbalance issues: majority classes are over‑predicted while minority classes, especially EC6, have zero recall despite reasonable ROC‑AUC. The authors recommend target‑specific threshold tuning and conformal calibration as post‑processing safeguards to expose and correct these hidden errors.

By Bilal Ahmad, Rajed Mehmood