How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
arXiv:2608. 12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets.
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint.
arXiv:2608. 12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets.
arXiv:2607. 28437v1 Announce Type: new Abstract: Molecular optimization is commonly performed under a limited oracle budget, which makes deciding what to evaluate as important as deciding what to generate.
arXiv:2606. 02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural constraints.
arXiv:2505. 20346v3 Announce Type: replace-cross Abstract: Function-guided protein design is a crucial task with significant applications in drug discovery and enzyme engineering.
arXiv:2606. 31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design.
arXiv:2602.00663v3 Announce Type: replace Abstract: Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical industr...
The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.
arXiv:2604.07669v3 Announce Type: replace-cross Abstract: Synthesizable molecular optimization seeks to improve target properties while ensuring that molecular modifications follow feasible synthetic...
arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.
arXiv:2609.08059v1 Announce Type: cross Abstract: Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for industrial biocatalysis, acid-related biopro...
The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.
Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good.