Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Training-Free Generation of Protein Sequences from Small Family Alignments via Stochastic Attention

About

Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples. Yet most protein families are small: the median Pfam seed alignment contains only 22 sequences, a regime where learned models overfit or collapse. We propose \emph{stochastic attention} (SA), a training-free sampler that treats the modern Hopfield energy over stored sequences as a Boltzmann distribution and draws samples via Langevin dynamics. The score function is the residual of a single softmax attention operation, eliminating the need for a trained score network, pretraining data, or graphics processing units (GPUs). Across eight Pfam families spanning 37 to 420 sequences and 23 to 262 residues, SA generates sequences with low composition divergence, novelty, and structural plausibility supported by ESMFold and AlphaFold2. Compared with profile hidden Markov models (HMMs), EvoDiff, and the multiple sequence alignment (MSA) Transformer, SA is the only tested method to simultaneously achieve low composition divergence, genuine novelty, and sequence identity within each family's nearest-neighbor identity range; the others drift outside this range or produce near-copies. The critical inverse temperature is predicted from principal component analysis (PCA) dimensionality alone, enabling fully automatic operation from a seed alignment. In two domains with deep mutational scanning data, SA-generated substitutions are enriched for experimentally tolerated mutations beyond a position-matched null, and an independent language model (ESM2-650M) scores them within the natural range. Stochastic attention thus opens training-free sequence generation to the long tail of protein families too small for deep learning.

Jeffrey D. Varner• 2026

Related benchmarks

TaskDatasetResultRank
Protein Sequence Pseudo-Perplexity EvaluationPfam Protein Families RRM, SH3, WW, Kunitz, zf-C2H2, PDZ, Pkinase, Defensin_beta
Pseudo-Perplexity4.44
16
Pairwise mutual information preservationKunitz
Pearson Correlation (r)0.8
6
Pairwise mutual information preservationSH3
Pearson Correlation (r)0.66
6
Sequence GenerationPfam RRM family
KL Divergence (AA)0.06
5
Sequence GenerationPfam WW family
KL Divergence (AA)0.008
5
Sequence GenerationPfam Kunitz family
KL Divergence (AA)0.013
5
Sequence GenerationPfam PDZ family
KL Divergence (AA)0.038
5
Sequence GenerationPfam Pkinase family
KL Divergence (AA)0.035
5
Pairwise mutual information preservationWW
Pearson Correlation (r)0.71
5
Pairwise mutual information preservationzf-C2H2
Pearson Correlation (r)0.92
5
Showing 10 of 19 rows

Other info

Follow for update