arXiv Machine Learning By Jonathan Legrand (IMB, MONC), Aguirre Mimoun (CHU Bordeaux), Baudouin Denis de Senneville (IMB, MONC), Audrey Bidet (CHU Bordeaux), Pierre-Yves Dumas (CHU Bordeaux, Inserm U1312 - BRIC), Christ\`ele Etchegaray (MONC, IMB)

Interpretable Multi-Instance Learning Enables Early Prediction of Key Molecular Alterations from Routine Flow Cytometry in Acute Myeloid Leukemia

Read the original on arXiv Machine Learning →

An interpretable multi‑instance learning classifier based on a decision tree was developed to predict NPM1 and FLT3‑ITD mutations in acute myeloid leukemia using routine flow cytometry data. In cross‑validation on 197 patients, the model achieved AUROCs of 0.96 for NPM1 and 0.86 for FLT3‑ITD, outperforming a clinical baseline and matching deep learning methods. On an independent cohort of 161 patients, it maintained high performance with AUROCs of 0.90 and 0.82, and positive predictive values of 0.87 and 0.68, while cell‑level interpretation recovered known immunophenotypic signatures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 12

Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

arXiv:2608. 10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios.

By Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
arXiv Machine Learning
Aug 11

TRAPS: Treatment-Assignment Prediction via Pathway-informed Stratification

arXiv:2606. 09898v2 Announce Type: replace Abstract: Cancer treatment involves decisions across multiple clinical outcomes, yet pathway-informed deep learning models are typically evaluated in isolation, making their relative benefits unclear.

By Sujoy Banik, Sayantan Chakraborty, Boishakhi Das Toma, Zainab Ghafoor, Ushashi Bhattacharjee, Koushik Howlader, Tirtho Roy
arXiv Computer Vision
Aug 26

EMFE: A lightweight, explainable machine learning framework for malaria cell classification

EMFE (Efficient Mathematical Feature Extraction) is a lightweight, explainable machine‑learning framework that classifies single red‑blood‑cell images as parasitized or uninfected using five engineered features: Gray World color normalization, adaptive green‑channel thresholding, morphological spot detection, and classical classifiers. On the NIH LHNCBC malaria dataset (27,558 images from 200 patients), a tuned Random Forest achieved 94.6% pooled out‑of‑fold accuracy, 94.3% on a 40‑patient holdout, and outperformed deep‑learning baselines in an accuracy‑efficiency trade‑off. Ablation studies, synthetic perturbations, and explainability analyses identified spot saturation as the dominant discriminative feature and quantified the framework’s failure modes and patient‑level performance.

By Md Abdullah Al Kafi, Walayat Hussain, Mousumi Karmakar, Sumit Kumar Banshal, Ahmed Al Marouf
arXiv Computer Vision
Aug 27

Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?

The study evaluates 15 frozen hematology foundation-model embeddings across four single‑cell acquisition domains, finding that while in‑domain accuracy is near‑saturated (macro‑F1 0.98–0.997), cross‑dataset performance drops dramatically (34–72%) and model rankings shift. Probe‑dependent rank transfer is observed, with 1‑NN retrieval more stable than linear heads, yet neither reliably predicts target robustness. Calibration deteriorates off‑domain (ECE rises from 0.004 to 0.35), and exposure to internal cohorts confounds shift analysis; a training‑free pseudo‑label‑balanced feature normalization (CBR) modestly improves target‑prior robustness and calibration. whyItMatters":"The findings highlight that frozen hematology foundation models, though accurate in‑domain, may fail under realistic scanner, site, and class‑prior shifts, underscoring the need for comprehensive audits of accuracy, calibration, exposure, and robustness before clinical deployment."

By Jai Kumar Sharma, Peeyush Tapadiya