arXiv AI

Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation

arXiv:2606. 14823v1 Announce Type: cross Abstract: Genetic evidence is enriched among approved drug targets: in an observational analysis of 26,278 target-disease pairs from Open Targets and ChEMBL, targets with any genetic association had a 3.

arXiv Machine Learning
Jun 10

OncoTraj: a public benchmark for longitudinal resistance prediction in EGFR-mutant non-small-cell lung cancer on osimertinib

arXiv:2606. 11144v1 Announce Type: new Abstract: Resistance to first-line osimertinib in EGFR-mutant non-small-cell lung cancer (NSCLC) is the canonical example of predictable clonal evolution under therapeutic pressure, yet no public benchmark exists for training or evaluating computational models on the corresponding longitudinal patient trajectories.

By Abhijoy Sarkar, Aarchi Singh Thakur
arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv AI
6d ago

FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification

FedHisto-PAST v2 is a parameter‑efficient, stain‑aware federated learning framework for cross‑site lung histopathology classification, combining a frozen HIBOU‑B foundation model with techniques such as paired‑view prediction, feature consistency, prototype learning, and adaptive aggregation. In a five‑client, non‑IID simulation and an exploratory LungHist700 cohort, the method achieved a Macro‑F1 of 0.7286 and a balanced accuracy of 0.7305, with the prediction‑level consistency component providing the most clear independent benefit. The framework updated only about 1.25% of the model parameters, demonstrating efficient adaptation while acknowledging limitations in privacy guarantees and clinical validation.

By Muhammad Muhtasim Shahriar, M. M. Golam Hafiz, Saad Aloteibi, Mohammad Ali Moni
Hugging Face Trending Papers
Aug 11

RadFusion: Towards Threshold-Controllable Radiology Report Generation

Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions.

arXiv AI
2d ago

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

OpenMTB‑Audit is an open‑source benchmark that tests large language models on 500 synthetic non‑small cell lung cancer cases, covering five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. The study found that all eight tested LLMs over‑refused Partially Supported recommendations, collapsing labels to achieve high safety scores. A deterministic seven‑module framework, MTB‑AuditAgent, was introduced to reduce over‑refusal to 6.7% and reach 91.2% accuracy, while an oncologist annotation study highlighted disagreement around the boundary between information sufficiency and treatment optimization.

By Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
arXiv Machine Learning
Sep 24

Benchmarking Active Spot Selection for Cost-Efficient Spatial Transcriptomics

The study benchmarks active spot selection methods against random sampling for spatial transcriptomics, focusing on cost‑efficient data acquisition. Using two public cohorts, the authors simulate multi‑round selection with uncertainty‑based (MC‑dropout, TOD) and diversity‑based (CoreSet, TypiClust) strategies, evaluating performance at 5%, 10%, 30%, and 50% of the spot pool. Results show that none of the active strategies consistently outperforms random sampling across all budgets or evaluation metrics, with performance varying by dataset and metric.

By Zheyu Zhu, Junchao Zhu, Fengbei Liu, Tianyuan Yao, Gelei Xu, John Cannon, Haichun Yang, Yuankai Huo, Mert R. Sabuncu, Ruining Deng
arXiv AI
Aug 12

RadFusion: Towards Threshold-Controllable Radiology Report Generation

arXiv:2608. 10505v1 Announce Type: new Abstract: Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content.

By Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz