arXiv Machine Learning

Testing when adaptive data acquisition can replace fixed measurement plans

arXiv:2607. 27651v2 Announce Type: replace Abstract: Learned rules select samples for follow-up measurements in high-throughput experiments.

arXiv Machine Learning
Sep 3

Held-out evidence resolves follow-up measurement decisions in biological screens

The paper introduces OPAL, a held‑out decision test that evaluates follow‑up measurement rules in biological screens by freezing a rule and assessing unnecessary measurement, coverage, and value after cost against pre‑defined archive‑specific criteria. Using a six‑rule Cell Painting battery, the authors show that a high‑value rule would re‑image 96.01% of the library with a 97.14% false‑activation upper bound, illustrating that predicted value alone cannot justify replacing a fixed plan. In development, a sparse Cell Painting rule reduced added‑well burden 18.2‑fold but had a false‑discovery bound above 35%, leading the fixed plan to remain; similar analyses for LINCS–LJP and CTRP highlighted the need for fallback strategies and the importance of separating optimization from evidence. "whyItMatters":"OPAL provides a systematic way to determine whether a new measurement strategy truly improves experimental efficiency without compromising data quality, as demonstrated across multiple biological screening datasets."

By Jia Bi, Samuel Pinilla, Chenyang Zhu
arXiv Computer Vision
Sep 4

SafeRestore: Detector-Relative Risk Certificates for Selective Industrial Image Restoration

SafeRestore introduces a framework for certifying when an industrial image restoration should be automatically returned to a detector or require human review. It ranks five restoration candidates using action‑specific fitted scores, selects a threshold gate on tuning data, and evaluates the gate on a separate certification sample with two one‑sided exact binomial bounds—one for evidence‑loss incidents and one for excess‑activation incidents. In a retrospective study of 4,591 Carinthia‑S images, the protocol demonstrates auditable risk‑coverage behavior, with varying pass rates across different policies and morphologies.

By Shaoliang Yang, Jun Wang
arXiv AI
Sep 12

When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents

The paper introduces an admission‑audit protocol for continual embodied agents, arguing that update admission should balance error control with retained learning opportunities within a fixed interaction budget. It critiques a range‑based confidence gate for failing to certify unchanged old‑task behavior, and proposes a paired‑binomial construction that reduces this burden when outcome disagreements are rare. Experiments on a one‑step pushing diagnostic show that fresh paired checks admit a significant portion of updates while the range‑based gate admits none, and a learned‑dynamics stress test helps distinguish model bias from feedback‑selection error.

By Qinzhen Ma, Ruihai Wu
arXiv AI
Aug 11

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.

By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
arXiv Computation and Language
Sep 11

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

The paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens across five phenotype categories, and builds on it to develop AssayLoop, a sequential experimental design framework that combines a transformer-based acquisition policy (AssayFormer) trained on historical data with LLM-derived biological priors. AssayLoop achieves a 5.67‑fold enrichment over random selection, recovering 27.7% of hits after testing only about 5% of the library, and outperforms existing adaptive-design methods and standalone LLMs. The authors also present AssayLLM, extending the approach directly to an LLM via task‑specific post‑training, and show that performance improves with more historical training data and transfers to unseen phenotype categories.

By Carl Edwards, Edward De Brouwer, Xiner Li, Namkyeong Lee, Ehsan Hajiramezanali, Anne Biton, Sara Mostafavi, Gabriele Scalia
arXiv AI
Sep 3

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.

By Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan
arXiv Machine Learning
Sep 15

Learning Source Acquisition Policies by Offline Planning

The paper introduces O-MPAC, an offline planning method for learning source acquisition policies under a limited budget. It transfers finite‑horizon risk‑cost targets from full training data into a shared source‑action scorer that re‑evaluates partial observations and source metadata after each query, applying a hard cost mask. Experiments show that O‑MPAC achieves high accuracy (0.965) in a routing task and outperforms several baselines on six real tasks, achieving the highest mean budget‑integrated accuracy on five of them.

By Ziqi Zhao, Run Xu, Qingjian Ni
arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv Machine Learning
Sep 10

Conditional Validity for Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.

By Melika Baghi