Active learning for affinity prediction of antibodies
About
The primary objective of most lead optimization campaigns is to enhance the binding affinity of ligands. For large molecules such as antibodies, identifying mutations that enhance antibody affinity is particularly challenging due to the combinatorial explosion of potential mutations. When the structure of the antibody-antigen complex is available, relative binding free energy (RBFE) methods can offer valuable insights into how different mutations will impact the potency and selectivity of a drug candidate, thereby reducing the reliance on costly and time-consuming wet-lab experiments. However, accurately simulating the physics of large molecules is computationally intensive. We present an active learning framework that iteratively proposes promising sequences for simulators to evaluate, thereby accelerating the search for improved binders. We explore different modeling approaches to identify the most effective surrogate model for this task, and evaluate our framework both using pre-computed pools of data and in a realistic full-loop setting.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Protein property prediction | 21 protein property datasets 48 data points (Cross-validation) | Spearman Correlation5.89 | 40 | |
| Protein property prediction | 21 protein landscapes Unseen mutations 96 data points | Spearman Correlation4.9 | 31 | |
| Protein property prediction | 21 protein landscapes 128 training points (extrapolation) | Spearman Correlation0.632 | 22 | |
| Protein property prediction | 21 protein landscapes Extrapolation 512 training points | Spearman Correlation0.739 | 22 | |
| Protein property prediction | 21 protein landscapes 1536 train points (cross-val) | Spearman Correlation0.846 | 22 | |
| Protein property prediction | 21 protein property datasets 128 data points (Extrapolation) | Spearman Correlation4.48 | 18 | |
| Protein property prediction | 21 protein property datasets 512 data points (Extrapolation) | Spearman Correlation4.38 | 18 | |
| Protein property prediction | 21 protein property datasets (Cross-validation) | Spearman Correlation0.846 | 9 | |
| Protein property prediction | Protein Property Datasets (Unseen mutations) | Spearman Correlation0.555 | 9 | |
| Protein property prediction | 21 protein landscapes (val) | Spearman Correlation3.76 | 9 |