arXiv AI

How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

arXiv:2608. 12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets.

arXiv Machine Learning
Aug 20

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.

By Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang
arXiv Machine Learning
Jun 9

Constraint-Aware Optimization for Robust Protein Stability Prediction

arXiv:2606. 08100v1 Announce Type: new Abstract: Multimodal $\Delta\Delta G$ predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on the Megascale dataset but exhibit limited robustness on out-of-distribution (OOD) proteins, persistent forward-reverse bias on paired-mutation benchmarks, and under-representation of rare stabilizing mutations.

By A Shivram, Aneesh S. Chivukula, Manik Gupta, Sourav Chowdhury
Hugging Face Trending Papers
Jul 29

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good.

arXiv AI
Aug 24

Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility

The study introduces Malaria-Instruct, a curated instruction-following dataset for malaria virtual screening, and evaluates five open-source large language models (Gemma-2, TxGemma, and LlaSMol-Mistral) against classical machine learning baselines and proprietary models. Fine‑tuned LLMs outperform all baselines, with TxGemma-9B achieving the highest ROC‑AUC (0.731 ± 0.005) and LlaSMol-Mistral-7B delivering the best enrichment factor (EF@1% ≈ 4.99). The results demonstrate that domain‑specific fine‑tuning and chemistry‑aware pretraining are essential for reliable discrimination, positioning fine‑tuned open‑source LLMs as a resource‑efficient alternative for antimalarial virtual screening.

By Marvellous O. Ajala (Magami Open Sciences Initiative), Zainab Ashimiyu-Abdusalam (Magami Open Sciences Initiative), Comfort Adesina (Magami Open Sciences Initiative)
Hugging Face Trending Papers
Aug 19

Off-Manifold Collapse in Guided Protein Language Models

The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.