The article introduces a new dataset for hepatocellular carcinoma (HCC) prediction, comprising 770 patient samples with 11,150 gene expression levels each, categorized into five classes from normal tissue to various HCC stages. The dataset was constructed using XGBoost and semi‑supervised learning on three existing genomic biomarker datasets, leveraging their labels to generate new annotations. The resulting XGBoost model achieved a 96.5% classification accuracy during the semi‑supervised training process.
By Ahmed Ammar Kubba, Manar Abu Talib, Jibran Sualeh Muhammad, Ali Bou Nassif, Abdalla Sayed Mohamed, Darko Castven, Jens U. Marquardt
arXiv:2408.07988v3 Announce Type: replace
Abstract: Despite significant research efforts and advancements, cancer remains a leading cause of mortality. Early cancer prediction has become a crucial fo...
By Samta Rani, Tanvir Ahmad, Sarfaraz Masood, Chandni Saxena
The paper introduces AMPLIFAI, the first public dataset of multiphase abdominal CT scans annotated with LI-RADS categories and segmented for three key LI-RADS features: arterial phase hyperenhancement, washout, and enhancing capsule. It outlines the dataset’s composition, curation process, and annotation pipeline following the Datasheets for Datasets format to promote transparency and reproducibility. The dataset aims to support the development of AI models for automated hepatocellular carcinoma diagnosis using the biopsy‑free, imaging‑based LI‑RADs framework.
By Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi, Jana Delfino, James Tonascia, Jade Wong-You-Cheong, Barton Lane, Joseph Chirico, Jeffrey D. Hirsch, Ang Li, Heng Huang, Florence X. Doo
arXiv:2606. 26561v1 Announce Type: new Abstract: Hepatitis C is a liver infection caused by a virus, which results in mild to severe inflammation of the liver.
By Abrar Alotaibi, Lujain Alnajrani, Nawal Alsheikh, Alhatoon Alanazy, Salam Alshammasi, Meshael Almusairii, Shoog Alrassan, Aisha Alansari
arXiv:2607. 03253v1 Announce Type: cross Abstract: As hematoxylin & eosin (H&E) staining constitutes the primary entry point in routine diagnostic workflows, computer-aided diagnosis from whole-slide H&E images is of particular clinical relevance.
By Ivica Kopriva, Dario Sitnik, Arijana Pacic, Karolina Krstanac, Irena Veliki Dalic, Marijana Popovic Hadzija
Hepatitis C is a liver infection caused by a virus, which results in mild to severe inflammation of the liver. Over many years, hepatitis C gradually damages the liver, often leading to permanent scarring, known as cirrhosis.
arXiv:2606. 09860v1 Announce Type: cross Abstract: Non-alcoholic fatty liver disease (NAFLD) affects roughly 25% of global adults, posing substantial hepatic and cardiovascular risks.
By Xinze Zhang
arXiv:2606. 04453v1 Announce Type: cross Abstract: Radiomics enables extraction of quantitative imaging biomarkers from medical images and has become an important tool for computer-aided cancer diagnosis.
By Hina Shakir, Mohammad Mohatram, Javeed Hussain, Syed Rizwan Ali, Muhammad Irfan Memon
arXiv:2607. 03466v1 Announce Type: cross Abstract: This study aims to predict Tumor, Node, and Metastasis (TNM) stage labels independently, with the Cancer Genome Atlas (TCGA) pathology report as the sixth shared task of SMM4H-HeaRD 2026.
By Joseph Itopa Abubakar, Jorge Jarme, Favour Igwezeke, Mary Adewunmi
arXiv:2601. 00175v2 Announce Type: replace Abstract: Objective: Develop and evaluate machine learning (ML) models for predicting incident liver cirrhosis (LC) one and two years prior to diagnosis using routinely collected electronic health record (EHR) data and benchmark their performance against the FIB-4 and APRI clinical scores.
By Zhuqi Miao, Ahmed G Qasem, Sujan Ravi, Jason T. Cheng, Abdulaziz Ahmed, Courtney W. Houchen, Sumayah Abed, Dilorom Azimdjanovna Zuparova, Abdulaziz Ahmed
arXiv:2606. 09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once.
By Sujoy Banik, Sayantan Chakraborty, Boishakhi Das Toma, Zainab Ghafoor, Ushashi Bhattacharjee, Koushik Howlader, Tirtho Roy
The study evaluates the use of default decision thresholds (t=0.50) in multi‑label enzyme commission (EC) number prediction across 14,096 compounds and six EC classes. It finds a high mean accuracy of 77.16% but low macro F1 (0.3976) and macro recall (0.3872), indicating severe class‑imbalance issues: majority classes are over‑predicted while minority classes, especially EC6, have zero recall despite reasonable ROC‑AUC. The authors recommend target‑specific threshold tuning and conformal calibration as post‑processing safeguards to expose and correct these hidden errors.
By Bilal Ahmad, Rajed Mehmood