arXiv AI

Analysis of Prompt Engineering for Drug Toxicity Prediction

The paper investigates how the phrasing of prompts affects large language models (LLMs) in predicting drug toxicity. By varying job role, prompt structure, and rule interpretation, the authors found that natural variability in LLM outputs outweighs fine‑tuning of prompts. However, incorporating chemoinformatic code to extract features significantly improved model performance, suggesting that prompt engineering alone is insufficient for reliable toxicity prediction.

arXiv Computation and Language
Sep 3

HarmReduction: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

The paper introduces HarmReduction, a benchmark for evaluating large language models (LLMs) on their ability to provide accurate and safe harm reduction information to people who use drugs (PWUD). The benchmark, HR-Basic, contains 2,160 question‑answer‑evidence pairs covering safety boundary checks, quantitative value provision, and polysubstance risk inference. Experiments show that even state‑of‑the‑art LLMs struggle with accuracy and can pose severe safety risks, underscoring the need for a dedicated evaluation framework.

By Kaixuan Wang, Chenxin Diao, Jason T. Jacques, Zhongliang Guo, Shuai Zhao
arXiv Computation and Language
Sep 23

Learning Diagnostic Reasoning for Decision Support in Toxicology

The paper introduces DeToxR, a reinforcement‑learning‑enhanced large language model designed to support decision making in acute toxicology cases. It fuses unstructured narratives from paramedics and patients with structured vital‑sign data to predict co‑ingested substances across 14 classes. In preliminary validation, DeToxR outperforms baseline models, achieving higher micro‑F1 and recall scores for poison identification.

By Nico Oberl\"ander, David Bani-Harouni, Tobias Zellner, Nassir Navab, Florian Eyer, Matthias Keicher
arXiv Machine Learning
Aug 28

Predicting Quantifiability from Primary Screens to Prioritize Dose-Response Profiling

The paper introduces a framework to predict whether a compound’s potency can be quantified in dose‑response profiling, treating quantifiability as a separate triage goal from biological activity. It shows that features from low‑cost primary screens, rather than molecular structure, strongly predict quantifiability, and that this prediction holds across new chemical scaffolds and assay families. The authors argue that incorporating quantifiability predictions can better allocate expensive dose‑response resources.

By Sean Lim
arXiv AI
Aug 28

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

The paper introduces the Latent Diagnostic Taxonomy, a framework that builds a dimensionality‑optimized classifier and a diagnostic tool to assess the trustworthiness of its confident predictions. It identifies a small set of influential prompts (latent support vectors) that reveal tokens which can change the classifier’s output, and uses these tokens to create a taxonomy that classifies prompts into safe, heuristic bias, heuristic override, or insufficient context categories. Applied to a prompt‑injection detection model, the framework shows that about 77% of confident decisions are fragile to a single token, distinguishing between calibration failures and exploitable shortcuts, and offers remediation strategies for each taxonomy zone.

By Jaturong Kongmanee, Smile Thanapattheerakul