Conditioning Protein Generation via Hopfield Pattern Multiplicity
arXiv:2603. 20115v2 Announce Type: replace Abstract: Small protein-family alignments often contain a subset of interest but not enough labeled data to train a conditional generator.
arXiv:2603. 14717v2 Announce Type: replace Abstract: Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples.
arXiv:2603. 20115v2 Announce Type: replace Abstract: Small protein-family alignments often contain a subset of interest but not enough labeled data to train a conditional generator.
arXiv:2608. 00697v1 Announce Type: cross Abstract: Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design.
arXiv:2603.29529v2 Announce Type: replace-cross Abstract: Using a statistical mechanics framework, we investigate the parameter space of transformer models trained on protein sequence data. We sample...
arXiv:2609.37675v1 Announce Type: new Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
arXiv:2507. 08920v4 Announce Type: replace-cross Abstract: We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, encompassing pretraining scaling laws, emergent capability analysis, in-context learning mechanism, and test-time scaling algorithm.
The study evaluates 4‑bit quantization and low‑rank adapter fine‑tuning (QLoRA) on several large protein language models, finding that many model‑task pairs retain over 90% of full fine‑tuning performance while achieving up to 90% GPU memory savings. QLoRA preserves early‑layer representations and induces task‑specific changes in later layers, closely resembling full fine‑tuning with smaller representational shifts. For generative models, 4‑bit quantization largely maintains structural and sequence‑level properties, though token‑level analysis reveals model‑dependent changes in autoregressive output distributions.
arXiv:2605. 00182v3 Announce Type: replace Abstract: Proteins are shaped by gradual evolution under biophysical and functional constraints.
SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.
arXiv:2602. 17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature".
arXiv:2606. 27361v1 Announce Type: cross Abstract: Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics.
arXiv:2510.03095v4 Announce Type: replace Abstract: Diffusion- and flow-based generative models have recently demonstrated strong performance in protein backbone generation tasks, offering unpreceden...
MT-ProtBERT is a multi‑task extension of ProtBERT designed for classifying intrinsically disordered proteins (IDPs) in low‑data settings. It combines Dynamic Window Masking, a Multi‑Scale 1D Convolutional classifier, and auxiliary biochemistry‑informed objectives to jointly optimize masked language modeling and domain‑specific tasks. In experiments on phosphorylation site prediction and protein compaction prediction, MT‑ProtBERT outperforms the RNN‑based IDP model PARROT across all limited‑data tasks.