arXiv Machine Learning By Jia Bi, Samuel Pinilla, Chenyang Zhu

Held-out evidence resolves follow-up measurement decisions in biological screens

Read the original on arXiv Machine Learning →

The paper introduces OPAL, a held‑out decision test that evaluates follow‑up measurement rules in biological screens by freezing a rule and assessing unnecessary measurement, coverage, and value after cost against pre‑defined archive‑specific criteria. Using a six‑rule Cell Painting battery, the authors show that a high‑value rule would re‑image 96.01% of the library with a 97.14% false‑activation upper bound, illustrating that predicted value alone cannot justify replacing a fixed plan. In development, a sparse Cell Painting rule reduced added‑well burden 18.2‑fold but had a false‑discovery bound above 35%, leading the fixed plan to remain; similar analyses for LINCS–LJP and CTRP highlighted the need for fallback strategies and the importance of separating optimization from evidence. "whyItMatters":"OPAL provides a systematic way to determine whether a new measurement strategy truly improves experimental efficiency without compromising data quality, as demonstrated across multiple biological screening datasets."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 11

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

The paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens across five phenotype categories, and builds on it to develop AssayLoop, a sequential experimental design framework that combines a transformer-based acquisition policy (AssayFormer) trained on historical data with LLM-derived biological priors. AssayLoop achieves a 5.67‑fold enrichment over random selection, recovering 27.7% of hits after testing only about 5% of the library, and outperforms existing adaptive-design methods and standalone LLMs. The authors also present AssayLLM, extending the approach directly to an LLM via task‑specific post‑training, and show that performance improves with more historical training data and transfers to unseen phenotype categories.

By Carl Edwards, Edward De Brouwer, Xiner Li, Namkyeong Lee, Ehsan Hajiramezanali, Anne Biton, Sara Mostafavi, Gabriele Scalia
arXiv Machine Learning
5d ago

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv AI
Aug 25

Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN

The paper investigates how to allocate limited amyloid PET scans in Alzheimer’s research by comparing simple target‑aligned validation strategies to more complex uncertainty‑based sampling. Using the A4/LEARN PET archive, it shows that for the APOE4 carrier versus non‑carrier contrast, balancing scans by APOE4 status nearly matches the performance of a target‑specific scoring approach, while generic uncertainty sampling performs worse. For other analyses, such as age‑slope or cutoff‑indexed PET positivity, target‑specific scoring yields greater gains, underscoring that measurement allocation should align with the specific claim being validated.

By Eliuvish Han Cui