Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
arXiv:2506. 10912v4 Announce Type: replace Abstract: Toxicity remains a leading cause of early-stage drug development failure.
The paper investigates how the phrasing of prompts affects large language models (LLMs) in predicting drug toxicity. By varying job role, prompt structure, and rule interpretation, the authors found that natural variability in LLM outputs outweighs fine‑tuning of prompts. However, incorporating chemoinformatic code to extract features significantly improved model performance, suggesting that prompt engineering alone is insufficient for reliable toxicity prediction.
arXiv:2506. 10912v4 Announce Type: replace Abstract: Toxicity remains a leading cause of early-stage drug development failure.
The paper introduces HarmReduction, a benchmark for evaluating large language models (LLMs) on their ability to provide accurate and safe harm reduction information to people who use drugs (PWUD). The benchmark, HR-Basic, contains 2,160 question‑answer‑evidence pairs covering safety boundary checks, quantitative value provision, and polysubstance risk inference. Experiments show that even state‑of‑the‑art LLMs struggle with accuracy and can pose severe safety risks, underscoring the need for a dedicated evaluation framework.
The paper introduces DeToxR, a reinforcement‑learning‑enhanced large language model designed to support decision making in acute toxicology cases. It fuses unstructured narratives from paramedics and patients with structured vital‑sign data to predict co‑ingested substances across 14 classes. In preliminary validation, DeToxR outperforms baseline models, achieving higher micro‑F1 and recall scores for poison identification.
arXiv:2607. 12886v1 Announce Type: new Abstract: Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields.
The paper introduces a framework to predict whether a compound’s potency can be quantified in dose‑response profiling, treating quantifiability as a separate triage goal from biological activity. It shows that features from low‑cost primary screens, rather than molecular structure, strongly predict quantifiability, and that this prediction holds across new chemical scaffolds and assay families. The authors argue that incorporating quantifiability predictions can better allocate expensive dose‑response resources.
arXiv:2609.21859v1 Announce Type: new Abstract: Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely o...
arXiv:2606. 14149v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in healthcare settings, yet their tendency to hallucinate poses risks when clinical decisions are involved.
arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.
arXiv:2408. 13378v5 Announce Type: replace Abstract: Workflows in drug-target interaction (DTI) assessment require integrating heterogeneous data from predictive models, curated resources, and observations from experimental literature.
This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be t...
arXiv:2606. 17113v1 Announce Type: new Abstract: Distinguishing causal adverse drug events (ADEs) from spurious correlations remains a central challenge in pharmacovigilance.
The paper introduces the Latent Diagnostic Taxonomy, a framework that builds a dimensionality‑optimized classifier and a diagnostic tool to assess the trustworthiness of its confident predictions. It identifies a small set of influential prompts (latent support vectors) that reveal tokens which can change the classifier’s output, and uses these tokens to create a taxonomy that classifies prompts into safe, heuristic bias, heuristic override, or insufficient context categories. Applied to a prompt‑injection detection model, the framework shows that about 77% of confident decisions are fragile to a single token, distinguishing between calibration failures and exploitable shortcuts, and offers remediation strategies for each taxonomy zone.