PAPER KEY: X66YQXSI
TITLE: AFsample2: Predicting multiple conformations and ensembles with AlphaFold2
AUTHORS: Wallner, Björn; Kalakoti, Yogesh

bioRxiv preprint doi: https://doi.org/10.1101/2024.05.28.596195; this version posted June 2, 2024. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC 4.0 International license.

AFsample2: Predicting multiple conformations and ensembles with AlphaFold2

Yogesh Kalakotia and Björn Wallnera
aDivision of Bioinformatics, Department of Physics, Chemistry and Biology, Linköping University, 581 83 Linköping, Sweden

1 Abstract

48 driven by the type and extent of structural dynamics asso-

2 Understanding protein dynamics and conformational states car- 49 ciated with the protein system. Conventional experimen3 ries profound scientiﬁc and practical implications for several ar- 50 tal structural biology methods such as X-ray crystallogra4 eas of research, ranging from a general understanding of biolog- 51 phy and cryogenic electron microscopy (cryo-EM) can pro5 ical processes at the molecular level to a detailed understanding 52 vide a few highly accurate snapshots of the overall confor-
6 of disease mechanisms, which in turn can open up new avenues 53 mational ensemble of the protein system (4–6). However,

7 in drug development. Multiple solutions have been recently de- 54 these snapshots only represent a fraction of possible states

8 veloped to widen the conformational landscape of predictions 55 and have to be supplemented by molecular dynamics or other

9 made by Alphafold2 (AF2). Here, we introduce AFsample2, a 56 similar solutions to infer molecular mechanisms. Further-

10

method employing random MSA column masking to reduce the 57

more, computational costs related to MD at biologically rel-

11

inﬂuence of co-evolutionary signals to enhance the structural di58

evant timescales are not viable in practice. Other experimen-

12 13 14

versity of models generated by the AF2 neural network. AFsample2 improves the prediction of alternative states for a broad 59 range of proteins, yielding high-quality end states and diverse 60

tal methods such as Nuclear Magnetic Resonance (NMR) could potentially proﬁle the dynamic nature of the protein

15 conformational ensembles. In the data set of open-closed con- 61 molecule, but are limited by scale (7).

16 formations (OC23), alternate state models improved in 17 out of 62

Recent advancements in in-silico protein structure de-

17 23 cases without compromising the generation of the preferred 63 termination have largely been an outcome of intelligent data

18

state. Consistent results were observed in 16 membrane protein 64

processing and generative artiﬁcial intelligence (AI). Meth-

19

transporters, with improvements in 12 out of 16 targets. TM65

ods like AlphaFold2 (8) (AF2) and RosettaFold (9) have

20 21 22

score improvements to experimental end states were substantial, sometimes exceeding 50%, elevating mediocre scores from 66 0.58 to nearly perfect 0.98. Furthermore, AFsample2 increased 67

demonstrated exceptional levels of success in determining accurate protein structure from evolutionary sequence infor-

23 the diversity of intermediate conformations by 70% compared 68 mation provided as Multiple Sequence Alignments (MSAs).

24 to the standard AF2 system, producing highly conﬁdent mod- 69 However, the default versions of these workﬂows are trained

25 els, that could potentially be on-path between the two states. In 70 to estimate a single high-conﬁdent model of the structure

26 addition, we also propose a way of selecting the end-states in 71 of a protein. This is a limitation since the entire confor-

27 generated model ensembles. These solutions could potentially 72 mational landscape has to be considered in order to get in-

28 enhance the generation and identiﬁcation of alternative protein 73 sights into the mechanistic basis of protein function. There-

29 conformations, thereby providing a more comprehensive under- 74 fore, an ideal sequence-to-structure prediction system should

30

standing of protein function and dynamics. Future work will 75

have the ability to model the entire conformational ensemble

31

focus on validating the accuracy of these intermediate confor76

for a given protein, identify states, and trace physically vi-

32 33

mations and exploring their relevance to functional transitions

in proteins.

77

78

able paths in estimated ensembles. Our recently developed AFsample method (10) captured different conformations of

34

Sampling | Alphafold | Conformations | Ensembles | Diversity

79 multimeric proteins by increasing the sampling rate and in-

35 Correspondence: bjorn.wallner@liu.se

80 troducing noise by enabling dropout layers at inference. The

81 method achieved state-of-the-art performance and was one of

36 Introduction

82 the top-ranked at CASP15 (11) (2022). Additional strategies 83 have also been proposed to induce conformational diversity

37 Proteins are the workhorses of life, serving as the build- 84 in AF2 predictions by subsampling the MSA, using shallow 38 ing blocks of cells and playing crucial roles in almost ev- 85 MSAs (12), in-silico alanine scan as in SPEACH_AF (13), or 39 ery biological process. They exhibit a wide range of func- 86 clustering the MSA as in AFcluster (14). All of these meth40 tions, including catalyzing biochemical reactions, providing 87 ods work by effectively reducing the information to AF2 to 41 structural reinforcement, and even acting as conduits in in- 88 allow the system to explore alternative solutions.

42 tracellular communication (1). Proteins adopt intricate three- 89

In this work, we present AFsample2, which employs

43 dimensional conﬁgurations, often existing within structural 90 random MSA column masking to diminish the constraints

44 ensembles that exhibit various states, collective movements, 91 exerted by co-evolutionary signals in MSA. Thereby increas-

45 and dynamic ﬂuctuations, all essential for executing their 92 ing the structural heterogeneity of models generated with the

46 function (2, 3). Processes such as folding, signal transduc- 93 AF2 neural network. AFsample2 was able to improve the

47 tion, enzyme catalysis, and molecular recognition are all 94 prediction of alternative states for a wide range of proteins.

Kalakoti et al. | bioRχiv | May 28, 2024 | 1–22

bioRxiv preprint doi: https://doi.org/10.1101/2024.05.28.596195; this version posted June 2, 2024. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC 4.0 International license.

Co-evolving residues

Retained

Contact Close conformation

PGSRADV

PGSRADV

HGGR ANR

HGGR ANR

K L QRMN A

XXXXXXX

L GQR AQA

XXXXXXX

Shallow MSA, MSA clustering and subsampling (Co-variance information retained)
Co-evolving residues

Masked
XX

d1
Open conformation d1<d2

PGSRADV

PGSRXDX

HGGR ANR

HGGR XNX

d2

Contact

K L QRMN A L GQR AQA

K LQRXNX L GQR XQX

Masked co-evolution

(a) Two strategies to alter the MSA, (top) MSA subsampling act on the rows of the MSAs, (bottom) column masking, break co-evolving residues and potentially contact networks.

(AvTerM-asgceordeoovferbe2s3t prmootdeielns) Model confidence

0.92

0.90

0.88

0.86

0.84

0.82

0.80

0.78

Best open Best close

0.76 00 05 10 15 20 25 30 35 40 50

MSA randomization (%)

90

85

80

75

70

Mean confidence

65

Best open confidence Best close confidence

00 05 10 15 20 25 30 35 40 50 MSA randomization (%)

(b) Average TM-scores and conﬁdences of best models for OC23 dataset with increasing MSA randomization

543322115000505050 543322115000505050
TM score (open conformation)
109988776655443322110505050505050505050500000000000000000000
TM score (close conformation)
109988776655443322110505050505050505050500000000000000000000

Best open models

Best close models

A0A07QQQQQQQAQOQABPPPPPPP5A5B7596903919730426739Q237DFZUXQEE8X61001213SR0WI9A4Y6RAV71515442SE9TWJMNNUR5682353894819ETTP82641933508381957412870

0.95 0.90 0.85 0.80 0.75 0.70 0.65 0.60

0.95 0.90 0.90 0.88 0.85 0.86 0.80 0.84 0.75 0.82 0.70 0.80 0.65 0.78
0.60

MSA randomisation (%)

Number of samples

%
00 05 10 15 20 25 30 35 40 50
Number of samples

(c) Per target best open and closed models for different MSA randomizations

(d) Highest TM-scores for open and close conformation with number of samples

Fig. 1. Overall summary and analysis of MSA randomization strategy in AFsample2. (a) A general outline of the modiﬁcations to the MSA for the AFsample2 pipeline (bottom). Traditional methods retain the information on co-variation, which in turn constrains the inference system to generate structures. AFsample2, on the other hand, remove those constraints by masking columns in the MSA to partially remove this co-variance information, leading to the generation of alternate conformations. (b) Effectiveness of the randomization strategy in terms of generating high-quality models and aggregate conﬁdence for both open and closed states. The results indicate 15% randomization to have the highest TM-scores in the OC23 dataset. (c) Optimal level of randomization on a per target protein. (d) Sampling more models increases the chances of generating better models, and is signiﬁcantly more potent with the proposed randomizations.

95 The improvement was quantiﬁed based on the ability of the 118 been presented in this study that enhance the capability of

96 inference system to generate high-quality end states and di- 119 MSA-based generative models in capturing the conforma-

97 verse conformational ensembles. The models for in particular 120 tional landscape of a given protein system.

98 the alternate state is improved for most of the cases (17/23)

99 in the open-closed data set (OC23) without sacriﬁcing per-
100 formance for the preferred state. The performance is main- 121 Results

101

tain on an additional set of 16 membrane protein transporters, 122

Method Development. The primary objective of this study

102

with the alternate state improved for 12/16 targets. The im123

was to improve the sampling of conformational states by

103

provement as measured by TM-score to experimental end 124

introducing more noise than simply turning on the dropout

104

states is sometimes massive with improvements over 50%, 125

layers at inference. In AFsample2, the noise is introduced

105

basically going from mediocre TM-scores of 0.58 to almost 126

by randomly masking columns in the MSA, with the ra-

106

perfect 0.98. However, the improvements are not only in the 127

tionale to break covariance constraints in the MSA, see

107

end-states but AFsample2 also improves the diversity by gen128

Fig 1a. By breaking covariance signals, the inference sys-

108

erating 70% more conformations in-between the end-states, 129

tem is allowed to explore and arrive at different solutions

109

when compared to the the vanilla AF2 system. While it re130

for the given protein, ultimately increasing the diversity of

110

mains to be demonstrated whether these intermediate confor131

the generated protein ensemble. A similar strategy for intro-

111

mations are accurate on-path representations between states, 132

ducing noise to MSAs has previously been attempted with

112

they are highly conﬁdent models generated by the AF2 infer133

SPEACH_AF (13), where a sliding window of alanines was

113 ence system.

134 introduced at speciﬁc columns in the MSA to break interact-

114

Furthermore, a novel strategy to identify conformational 135 ing residues. Although effective, this strategy was dependent

115 states from a pool of generated models without the aid of 136 on in-silico mutagenesis of the MSA, requiring prior knowl-

116 any experimental reference structures was also developed. 137 edge of the interacting residues. In their implementation,

117 In summary, signiﬁcant methodological improvements have 138 these residues were based on either prior structural informa-

2 | bioRχiv

Kalakoti et al. | AFsample2

bioRxiv preprint doi: https://doi.org/10.1101/2024.05.28.596195; this version posted June 2, 2024. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC 4.0 International license.

139 tion or/and contacts in generated models. AFsample2 does

Protein sequence

AlphaFold inference system

140 not have this limitation and provides a general solution even 141 for cases where such information is unavailable.
142 Effect of MSA masking on generated models. The amount

Genetic database search hhblits/MMseqs
Multiple sequence alignment (MSA)

RAW MSA
PAIRING

MSA track Randomized MSA masking
Pair representation track

Default AF2 architecture
Evoformer>Structure module -> recycles

143 of MSA masking, i.e., the fraction of randomized positions, 144 was observed to be the most important factor in the ability 145 of the inference system to generate alternate conformational 146 states. It was observed that increasing the MSA masking in147 creased the chances of generating end-state conformations

Templates

Reference-free State identification

Generated models (1000)
Rank by confidence

Screen 1: Confidence

Screen 2: Extremity Coverage

x1000
No References available ? Yes Ensemble diversity analysis

Model confidence

148 for a given protein. This trend is summarized in Fig 1b, 149 where MSA masking generates signiﬁcantly better models 150 compared to no masking (0%) for the alternate state (open in 151 these cases) across a set of diverse proteins with well-deﬁned 152 open and closed conformations (the OC23 set, see Methods). 153 The aggregate TM-score for the best alternate (open) confor-

Compute Similarity (TMscore) of top ranked model with all models

TM-score with best

Confidence threshold
TM-score with best

State2

Confidence and extremity-based screening

Identified states

State-guided ensemble analysis for
diversity plots and fill ratio

154 mation increases from 0.795 for no masking to 0.878 with 155 15% masking while showing a marginal improvement from

Fig. 2. Overall workﬂow of the AFsample2 pipeline. It starts by generating MSAs for a given protein sequence. This is followed by randomized MSA masking

156 0.89 to 0.90 for the closed conformation. Beyond 30% mask157 ing, performance drops ﬁrst for the open conformations and 158 subsequently for the closed conformations.

in a way such that a unique MSA proﬁle is fed into the system at every instance of the inference run. The generated ensemble is either passed to the diversity analysis workﬂow or the state-identiﬁcation workﬂow, depending on the availability of reference states.

159

In addition, it has previously been reported that the

160 model conﬁdence of AF2 predictions deteriorates with in- 195 ﬂecting the fact that AF2 has a preference for predicting the

161 creased sub-sampling (15). A similar trend was also observed 196 closed conformation in this case, leaving more room for im-

162 here, where the mean conﬁdence gradually decreased with 197 provement of the open conformation. The increasing trend

163 increasing MSA masking (Fig. 1b). The decrease is linear 198 is most pronounced for fewer samples, reﬂecting the switch

164 from 0% to 35% with a 2% drop in conﬁdence for every 5 199 from no sampling to actual sampling, but it is still increas-

165 percentage points of masking, followed by a rapid drop in 200 ing even up to 1000 samples, indicating that sampling more

166 model conﬁdence beyond 35% masking. Since MSA mask- 201 is always better. However, considering the trade-off between

167 ing essentially removes information, this trend is expected. 202 speed and performance, 1000 samples at 15% masking is a

168 However, it is important to realize that the decrease in model 203 reasonable default.

169 conﬁdence up to 20% masking is not coupled to lower qual170 ity models. It is most likely an effect that the mask itself 204 Overview of the AFsample2. Given a protein sequence, AF171 renders more uncertainty in the prediction, which in turn re- 205 sample2 follows a four-step process to generate diverse pro172 sults in lower model conﬁdence. Overall, 15% randomiza- 206 tein structures using a modiﬁed version of the AF2 inference 173 tion seems to perform marginally better than other settings. 207 system. It starts by (i) querying sequence databases to gen174 However, by analyzing the per target performance (Fig. 1c), 208 erate multiple sequence alignments (MSAs), (ii) Randomly 175 it can be seen that different levels of masking yield the best 209 masking MSA columns with a pre-deﬁned probability (e.g. 176 performance for different target proteins, e.g., 20% masking 210 15%), (iii) running inference on a uniquely masked MSA for 177 generates the best model for P40131, while 5% masking is 211 each model and lastly, (iv) depending on the availability of 178 optimal for P71147. The best TM-scores for each protein at 212 reference states, identifying state representatives with clus179 various masking levels can be visualized in Fig. S1. Even 213 tering, conﬁdence and extremity selection, followed by en180 though the exact magnitude of masking might differ between 214 semble analysis. A schematic representation of the workﬂow 181 targets, it is true that in all cases, masking is always better 215 is summarized in Fig. 2.

182 than no masking for the same level of sampling. For compar-

183 ative analysis and simplicity, AFsample2 using 15% masking 216 Comparing AFsample2 to AFvanilla, AFdropout and

184 was used for the downstream analysis.

217 AFcluster. The performance of AFsample2 was compared to

218 standard AF2 (AFvanilla), standard AF2 with dropout (AF-

185 The importance of sampling. It has been previously estab- 219 dropout), and AFcluster (14) by generating 1000 models for

186 lished that increased sampling improves the chances of gen- 220 each protein in the OC23 dataset. Fig. 3a shows the distri-

187 erating alternative conformations (10, 16). But the question 221 bution of TM-scores for the open and closed states for the

188 is, how much sampling is enough? To answer this ques- 222 models generated by the four methods. In terms of sampling

189 tion, we estimated the best TM-score for the open and closed 223 different states, the performance of AFvanilla and AFdropout

190 states, respectively, as the number of samples increased and 224 is almost identical; both generally have fairly narrow distribu-

191 for different levels of masking 0-50%, see Fig. 1d. Indeed, 225 tions centered on high TM-score for the closed state. In con-

192 generating more sampling increases the chances of generat- 226 trast, the distributions from AFsample2 cover a wider range,

193 ing higher-quality models for all levels of masking. The im- 227 and importantly, the distributions for the open state cover the

194 provement is larger for predicting the open conformation, re- 228 high TM-score region. This is unlike AFcluster, where even

Kalakoti et al. | AFsample2

bioRχiv | 3

bioRxiv preprint doi: https://doi.org/10.1101/2024.05.28.596195; this version posted June 2, 2024. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC 4.0 International license.

Target
A2RJ53 O76728 P31133 P00558 P40131 Q7DAU8 P21589 A0QTT2 Q5F9M1 Q9X6R4 Q18A65 Q9ERE7 P62495 A0A075Q0W3 P71447 P33284 Q9Z4N6 A6UVT1 B7IE18 B3EYN2 Q53W80 Q9SS90 Q9X9P9
Mean

BaselineTM between states
0.507 0.523 0.565 0.598 0.612 0.617 0.621 0.633 0.636 0.648 0.656 0.664 0.708 0.720 0.748 0.749 0.749 0.754 0.755 0.766 0.781 0.824 0.837
0.681

Similarity to o