Testing when adaptive data acquisition can replace fixed measurement plans
arXiv:2607. 27651v2 Announce Type: replace Abstract: Learned rules select samples for follow-up measurements in high-throughput experiments.
The paper introduces OPAL, a held‑out decision test that evaluates follow‑up measurement rules in biological screens by freezing a rule and assessing unnecessary measurement, coverage, and value after cost against pre‑defined archive‑specific criteria. Using a six‑rule Cell Painting battery, the authors show that a high‑value rule would re‑image 96.01% of the library with a 97.14% false‑activation upper bound, illustrating that predicted value alone cannot justify replacing a fixed plan. In development, a sparse Cell Painting rule reduced added‑well burden 18.2‑fold but had a false‑discovery bound above 35%, leading the fixed plan to remain; similar analyses for LINCS–LJP and CTRP highlighted the need for fallback strategies and the importance of separating optimization from evidence. "whyItMatters":"OPAL provides a systematic way to determine whether a new measurement strategy truly improves experimental efficiency without compromising data quality, as demonstrated across multiple biological screening datasets."
arXiv:2607. 27651v2 Announce Type: replace Abstract: Learned rules select samples for follow-up measurements in high-throughput experiments.
arXiv:2607. 27651v1 Announce Type: new Abstract: Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted.
The paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens across five phenotype categories, and builds on it to develop AssayLoop, a sequential experimental design framework that combines a transformer-based acquisition policy (AssayFormer) trained on historical data with LLM-derived biological priors. AssayLoop achieves a 5.67‑fold enrichment over random selection, recovering 27.7% of hits after testing only about 5% of the library, and outperforms existing adaptive-design methods and standalone LLMs. The authors also present AssayLLM, extending the approach directly to an LLM via task‑specific post‑training, and show that performance improves with more historical training data and transfers to unseen phenotype categories.
arXiv:2604. 11305v3 Announce Type: replace Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR).
The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.
The paper investigates how to allocate limited amyloid PET scans in Alzheimer’s research by comparing simple target‑aligned validation strategies to more complex uncertainty‑based sampling. Using the A4/LEARN PET archive, it shows that for the APOE4 carrier versus non‑carrier contrast, balancing scans by APOE4 status nearly matches the performance of a target‑specific scoring approach, while generic uncertainty sampling performs worse. For other analyses, such as age‑slope or cutoff‑indexed PET positivity, target‑specific scoring yields greater gains, underscoring that measurement allocation should align with the specific claim being validated.
arXiv:2609.21859v1 Announce Type: new Abstract: Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely o...
The paper introduces a framework to predict whether a compound’s potency can be quantified in dose‑response profiling, treating quantifiability as a separate triage goal from biological activity. It shows that features from low‑cost primary screens, rather than molecular structure, strongly predict quantifiability, and that this prediction holds across new chemical scaffolds and assay families. The authors argue that incorporating quantifiability predictions can better allocate expensive dose‑response resources.
arXiv:2607. 17345v1 Announce Type: new Abstract: Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options.
The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
arXiv:2607. 20950v1 Announce Type: new Abstract: BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably.