arXiv Machine Learning

Fixing a Model That Learned Worse Cancer Means Lower Risk: Monotonic Constraints in Bladder Cancer Recurrence Prediction

In a UK multicentre trial, an unconstrained XGBoost model incorrectly learned that higher tumour stage and carcinoma in situ predicted lower bladder cancer recurrence risk, a finding that conventional metrics such as discrimination, calibration, and SHAP failed to detect. The authors introduced a counterfactual direction test and a monotonic‑constraint framework, which removed the inversion without harming model performance and even outperformed established risk systems. The study demonstrates that such tests should be routine before deploying predictive models in clinical settings.

arXiv AI
Sep 18

FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity

The paper introduces FCA‑Guided Counterfactual (FCA‑CF) explanations for multi‑modal breast cancer diagnosis, leveraging a Formal Concept Analysis lattice as a hard structural constraint to search for counterfactuals. On the TCGA‑BRCA dataset, FCA‑CF achieves perfect validity (100% prediction flips), the lowest average feature changes (2.37), and competitive proximity (0.900) compared to four other methods. Ablation studies show the lattice constraint drives sparsity, while a greedy refinement phase further improves results.

By Abdullahi Isa, Souley Boukari, Muhammad Aliyu
Hugging Face Trending Papers
Sep 17

FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity

The paper introduces FCA‑Guided Counterfactual (FCA‑CF) explanations for multi‑modal breast cancer diagnosis, leveraging a Formal Concept Analysis lattice as a hard structural constraint to generate counterfactuals. On the TCGA‑BRCA dataset, FCA‑CF achieves perfect validity (100% prediction flips), the lowest average feature changes (2.37), and competitive proximity (0.900), outperforming four established counterfactual methods. Ablation studies show the lattice constraint and a greedy refinement phase are key to its sparsity and validity.

arXiv AI
Jul 10

Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

arXiv:2607. 08602v1 Announce Type: new Abstract: Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality.

By Peng Cui, Jitao Wang, Siyan Xue, Yao Huang, Haoming Xia, Dong Li, Dengxiang Liu, Weilin Wang, Liping Liu, Leida Zhang, Yunfu Cui, Tao Peng, Daolin Ji, Haitao Zhao, Wei Zhang, Xiaojuan Wang, Weijie Ma, Zongren Ding, Jinlong Li, Yuan Ding, Jiajing Zhao, Zhiyu Chen, Chengkun Yang, Ziyue Huang, Jiaqi Liu, Fusheng Liu, Yang Zhou, Xiaojuan Wang, Zhongquan Sun, Shiyun Bao, Xiaojun Wang, Ming Yang, Guangxin Li, Bin Shu, Yong Liao, Hongxuan Li, Yao Tang, Shizhong Yang, Yongyi Zeng, Yufeng Yuan, Yinpeng Dong, Jihui Hao, Jun Zhu, Jiahong Dong
arXiv AI
Sep 11

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

The paper introduces Cros, a risk‑constrained stopping layer for sequential clinical diagnosis agents that determines when to stop testing and make a diagnosis. Cros combines state‑wise error ranking, policy design on disjoint development splits, and exact tests of selective diagnostic error to provide finite‑sample guarantees. On a MIMIC‑derived abdominal‑pain benchmark, Cros achieves higher state‑error AUROC and lower selective error rates compared to baseline stopping methods, though its performance varies across development resplits.

By Yuexin Wu, Vasile Rus
arXiv Computer Vision
Sep 10

CHIMERA Challenge Task 2 and 3: Response Subtypes Classification and Progression Survival Prediction in Bladder Cancer Patients using Multimodal Datasets

arXiv:2609.09510v1 Announce Type: cross Abstract: High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratifi...

By Catherine Chia, Tongjie Wang, Robert Spaans, Maryam Mohammadlou, Farbod Khoraminia, J. Alberto Nakauma-Gonz\'alez, Adam Kowalewski, Parandzem Khachatryan, Domingos Oliveira, Khrystyna Faryna, CHIMERA Challenge Consortium, Marlies Wakkee, Sita Vermeulen, Tahlita Zuiverloon, Nadieh Khalili
arXiv AI
2d ago

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

OpenMTB‑Audit is an open‑source benchmark that tests large language models on 500 synthetic non‑small cell lung cancer cases, covering five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. The study found that all eight tested LLMs over‑refused Partially Supported recommendations, collapsing labels to achieve high safety scores. A deterministic seven‑module framework, MTB‑AuditAgent, was introduced to reduce over‑refusal to 6.7% and reach 91.2% accuracy, while an oncologist annotation study highlighted disagreement around the boundary between information sufficiency and treatment optimization.

By Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
arXiv AI
Aug 5

CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation

arXiv:2608. 03079v1 Announce Type: cross Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions.

By Ting Yin, Danning Li, Chen Shu, Xiaoxia Yao, Boyu Fu, Yujing Chang, Tianyu Shi, Mengna Feng, Jie Chen, Jing Fu, Xiuli Xiao, Tianlin Li, Mumin Shao, Jiaxin Bi, Wenchuan Zhang, Xiaoyan Wu, Xiao Han, Zhang Zhang, Yuhao Yi, Hong Bu
arXiv Machine Learning
Sep 2

TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening

The paper introduces TRUST, a threshold‑recalibrated training method that dynamically adjusts the dismissal threshold during training to penalize cancer‑positive images near the dismissal region. Evaluated on NLBS and RSNA datasets, TRUST achieved higher case‑level dismissal rates while maintaining 98% and 95% recall, outperforming a cross‑entropy baseline. External validation on RSNA→NLBS data confirmed improved dismissal rates at both recall targets, demonstrating the effectiveness of closed‑loop threshold‑aware training for selective dismissal in breast cancer screening.

By Parham Hajishafiezahramini, Matthew Hamilton, Edward Kendall, Gregory Doyle, Oscar Meruvia Pastor
arXiv Computer Vision
Aug 25

CHIMERA Challenge: Biochemical Recurrence Prediction in Prostate Cancer Patients using multimodal datasets

arXiv:2608.21497v1 Announce Type: cross Abstract: Biochemical recurrence (BCR), defined as any detectable prostate-specific antigen level after prostatectomy with confirmatory elevation, is widely us...

By Robert N. Spaans, Catherine Chia, Tongjie Wang, Adam Kowalewski, Parandzem Khachatryan, Domingos Oliveira, Khrystyna Faryna, Jean-Paul A. van Basten, Geert Litjens, Nadieh Khalili