arXiv Machine Learning

Certifying when decision-time information justifies adaptive experimentation

arXiv:2607. 27651v1 Announce Type: new Abstract: Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted.

arXiv Machine Learning
Sep 3

Held-out evidence resolves follow-up measurement decisions in biological screens

The paper introduces OPAL, a held‑out decision test that evaluates follow‑up measurement rules in biological screens by freezing a rule and assessing unnecessary measurement, coverage, and value after cost against pre‑defined archive‑specific criteria. Using a six‑rule Cell Painting battery, the authors show that a high‑value rule would re‑image 96.01% of the library with a 97.14% false‑activation upper bound, illustrating that predicted value alone cannot justify replacing a fixed plan. In development, a sparse Cell Painting rule reduced added‑well burden 18.2‑fold but had a false‑discovery bound above 35%, leading the fixed plan to remain; similar analyses for LINCS–LJP and CTRP highlighted the need for fallback strategies and the importance of separating optimization from evidence. "whyItMatters":"OPAL provides a systematic way to determine whether a new measurement strategy truly improves experimental efficiency without compromising data quality, as demonstrated across multiple biological screening datasets."

By Jia Bi, Samuel Pinilla, Chenyang Zhu
arXiv Machine Learning
Aug 28

Predicting Quantifiability from Primary Screens to Prioritize Dose-Response Profiling

The paper introduces a framework to predict whether a compound’s potency can be quantified in dose‑response profiling, treating quantifiability as a separate triage goal from biological activity. It shows that features from low‑cost primary screens, rather than molecular structure, strongly predict quantifiability, and that this prediction holds across new chemical scaffolds and assay families. The authors argue that incorporating quantifiability predictions can better allocate expensive dose‑response resources.

By Sean Lim
arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv Machine Learning
Aug 31

Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction

The study evaluates a frozen Geneformer representation for predicting CRISPRi perturbation effects under a tightly controlled, pre‑registered protocol. While the representation shows significant predictive power within the Virtual Cell Challenge dataset, it fails to transfer to external screens, with negative zero‑shot Spearman correlations. The analysis also reveals that the VCC endpoint is heavily influenced by sampling depth, as cell count alone explains most of the variance, indicating a sampling‑depth entanglement that could mask transfer failures in less controlled settings.

By Mehrdad Shoeibi, Niloofar Yousefi
arXiv Machine Learning
Jun 3

Fairness Definitions and Metrics in Deep Reinforcement Learning for Drug Discovery in Healthcare: A Rapid Evidence Review

arXiv:2606. 02902v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) is increasingly applied to de novo molecular design, but choices in data, rewards, and evaluation can yield uneven performance across disease areas and chemotypes.

By Esmaeil Shakeri, Ronnie de Souza Santos, Behrouz Far
arXiv AI
Jul 8

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

arXiv:2607. 05462v1 Announce Type: cross Abstract: As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse.

By Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman
arXiv Computation and Language
Sep 11

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

The paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens across five phenotype categories, and builds on it to develop AssayLoop, a sequential experimental design framework that combines a transformer-based acquisition policy (AssayFormer) trained on historical data with LLM-derived biological priors. AssayLoop achieves a 5.67‑fold enrichment over random selection, recovering 27.7% of hits after testing only about 5% of the library, and outperforms existing adaptive-design methods and standalone LLMs. The authors also present AssayLLM, extending the approach directly to an LLM via task‑specific post‑training, and show that performance improves with more historical training data and transfers to unseen phenotype categories.

By Carl Edwards, Edward De Brouwer, Xiner Li, Namkyeong Lee, Ehsan Hajiramezanali, Anne Biton, Sara Mostafavi, Gabriele Scalia
arXiv AI
Aug 11

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.

By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai