How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
arXiv:2608. 12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets.
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint.
arXiv:2608. 12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets.
arXiv:2607. 28437v1 Announce Type: new Abstract: Molecular optimization is commonly performed under a limited oracle budget, which makes deciding what to evaluate as important as deciding what to generate.
arXiv:2606. 02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural constraints.
arXiv:2505. 20346v3 Announce Type: replace-cross Abstract: Function-guided protein design is a crucial task with significant applications in drug discovery and enzyme engineering.
arXiv:2606. 31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design.
arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.
Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good.
arXiv:2607. 26397v1 Announce Type: cross Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification.
arXiv:2606. 09664v1 Announce Type: new Abstract: Bayesian optimization (BO) is a central tool for sample-efficient design, and latent-space Bayesian optimization (LSBO) extends it to structured objects such as molecules and proteins.
arXiv:2607. 26391v1 Announce Type: new Abstract: Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions.
arXiv:2512. 02328v2 Announce Type: replace-cross Abstract: Selecting an effective docking algorithm is highly context-dependent, and no single method performs reliably across structural, chemical, and protocol regimes.
arXiv:2602. 22822v3 Announce Type: replace Abstract: Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction still remain difficult to evaluate and deploy in practice.