1
|
Li C, Luo Y, Xie Y, Zhang Z, Liu Y, Zou L, Xiao F. Structural and functional prediction, evaluation, and validation in the post-sequencing era. Comput Struct Biotechnol J 2024; 23:446-451. [PMID: 38223342 PMCID: PMC10787220 DOI: 10.1016/j.csbj.2023.12.031] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/20/2023] [Revised: 12/20/2023] [Accepted: 12/22/2023] [Indexed: 01/16/2024] Open
Abstract
The surge of genome sequencing data has underlined substantial genetic variants of uncertain significance (VUS). The decryption of VUS discovered by sequencing poses a major challenge in the post-sequencing era. Although experimental assays have progressed in classifying VUS, only a tiny fraction of the human genes have been explored experimentally. Thus, it is urgently needed to generate state-of-the-art functional predictors of VUS in silico. Artificial intelligence (AI) is an invaluable tool to assist in the identification of VUS with high efficiency and accuracy. An increasing number of studies indicate that AI has brought an exciting acceleration in the interpretation of VUS, and our group has already used AI to develop protein structure-based prediction models. In this review, we provide an overview of the previous research on AI-based prediction of missense variants, and elucidate the challenges and opportunities for protein structure-based variant prediction in the post-sequencing era.
Collapse
Affiliation(s)
- Chang Li
- Clinical Biobank, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
- The Key Laboratory of Geriatrics, Beijing Institute of Geriatrics, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
| | - Yixuan Luo
- Beijing Normal University, Beijing, China
| | - Yibo Xie
- Information Center, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
| | - Zaifeng Zhang
- The Key Laboratory of Geriatrics, Beijing Institute of Geriatrics, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
| | - Ye Liu
- The Key Laboratory of Geriatrics, Beijing Institute of Geriatrics, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
| | - Lihui Zou
- The Key Laboratory of Geriatrics, Beijing Institute of Geriatrics, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
| | - Fei Xiao
- Clinical Biobank, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
- The Key Laboratory of Geriatrics, Beijing Institute of Geriatrics, Beijing Hospital, National Center of Gerontology, National Health Commission, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences, Beijing, China
- Beijing Normal University, Beijing, China
| |
Collapse
|
2
|
Zhou J, Huang M. Navigating the landscape of enzyme design: from molecular simulations to machine learning. Chem Soc Rev 2024. [PMID: 38990263 DOI: 10.1039/d4cs00196f] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 07/12/2024]
Abstract
Global environmental issues and sustainable development call for new technologies for fine chemical synthesis and waste valorization. Biocatalysis has attracted great attention as the alternative to the traditional organic synthesis. However, it is challenging to navigate the vast sequence space to identify those proteins with admirable biocatalytic functions. The recent development of deep-learning based structure prediction methods such as AlphaFold2 reinforced by different computational simulations or multiscale calculations has largely expanded the 3D structure databases and enabled structure-based design. While structure-based approaches shed light on site-specific enzyme engineering, they are not suitable for large-scale screening of potential biocatalysts. Effective utilization of big data using machine learning techniques opens up a new era for accelerated predictions. Here, we review the approaches and applications of structure-based and machine-learning guided enzyme design. We also provide our view on the challenges and perspectives on effectively employing enzyme design approaches integrating traditional molecular simulations and machine learning, and the importance of database construction and algorithm development in attaining predictive ML models to explore the sequence fitness landscape for the design of admirable biocatalysts.
Collapse
Affiliation(s)
- Jiahui Zhou
- School of Chemistry and Chemical Engineering, Queen's University, David Keir Building, Stranmillis Road, Belfast BT9 5AG, Northern Ireland, UK.
| | - Meilan Huang
- School of Chemistry and Chemical Engineering, Queen's University, David Keir Building, Stranmillis Road, Belfast BT9 5AG, Northern Ireland, UK.
| |
Collapse
|
3
|
Lombard V, Grudinin S, Laine E. Explaining Conformational Diversity in Protein Families through Molecular Motions. Sci Data 2024; 11:752. [PMID: 38987561 PMCID: PMC11237097 DOI: 10.1038/s41597-024-03524-5] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/31/2024] [Accepted: 06/14/2024] [Indexed: 07/12/2024] Open
Abstract
Proteins play a central role in biological processes, and understanding their conformational variability is crucial for unraveling their functional mechanisms. Recent advancements in high-throughput technologies have enhanced our knowledge of protein structures, yet predicting their multiple conformational states and motions remains challenging. This study introduces Dimensionality Analysis for protein Conformational Exploration (DANCE) for a systematic and comprehensive description of protein families conformational variability. DANCE accommodates both experimental and predicted structures. It is suitable for analysing anything from single proteins to superfamilies. Employing it, we clustered all experimentally resolved protein structures available in the Protein Data Bank into conformational collections and characterized them as sets of linear motions. The resource facilitates access and exploitation of the multiple states adopted by a protein and its homologs. Beyond descriptive analysis, we assessed classical dimensionality reduction techniques for sampling unseen states on a representative benchmark. This work improves our understanding of how proteins deform to perform their functions and opens ways to a standardised evaluation of methods designed to sample and generate protein conformations.
Collapse
Affiliation(s)
- Valentin Lombard
- Sorbonne Université, CNRS, IBPS, Laboratoire de Biologie Computationnelle et Quantitative (LCQB), 75005, Paris, France
| | - Sergei Grudinin
- Université Grenoble Alpes, CNRS, Grenoble INP, LJK, 38000, Grenoble, France.
| | - Elodie Laine
- Sorbonne Université, CNRS, IBPS, Laboratoire de Biologie Computationnelle et Quantitative (LCQB), 75005, Paris, France.
- Institut Universitaire de France (IUF), Paris, France.
| |
Collapse
|
4
|
Rahimzadeh F, Mohammad Khanli L, Salehpoor P, Golabi F, PourBahrami S. Unveiling the evolution of policies for enhancing protein structure predictions: A comprehensive analysis. Comput Biol Med 2024; 179:108815. [PMID: 38986287 DOI: 10.1016/j.compbiomed.2024.108815] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 03/04/2024] [Revised: 06/09/2024] [Accepted: 06/24/2024] [Indexed: 07/12/2024]
Abstract
Predicting protein structure is both fascinating and formidable, playing a crucial role in structure-based drug discovery and unraveling diseases with elusive origins. The Critical Assessment of Protein Structure Prediction (CASP) serves as a biannual battleground where global scientists converge to untangle the intricate relationships within amino acid chains. Two primary methods, Template-Based Modeling (TBM) and Template-Free (TF) strategies, dominate protein structure prediction. The trend has shifted towards Template-Free predictions due to their broader sequence coverage with fewer templates. The predictive process can be broadly classified into contact map, binned-distance, and real-valued distance predictions, each with distinctive strengths and limitations manifested through tailored loss functions. We have also introduced revolutionary end-to-end, and all-atom diffusion-based techniques that have transformed protein structure predictions. Recent advancements in deep learning techniques have significantly improved prediction accuracy, although the effectiveness is contingent upon the quality of input features derived from natural bio-physiochemical attributes and Multiple Sequence Alignments (MSA). Hence, the generation of high-quality MSA data holds paramount importance in harnessing informative input features for enhanced prediction outcomes. Remarkable successes have been achieved in protein structure prediction accuracy, however not enough for what structural knowledge was intended to, which implies need for development in some other aspects of the predictions. In this regard, scientists have opened other frontiers for protein structural prediction. The utilization of subsampling in multiple sequence alignment (MSA) and protein language modeling appears to be particularly promising in enhancing the accuracy and efficiency of predictions, ultimately aiding in drug discovery efforts. The exploration of predicting protein complex structure also opens up exciting opportunities to deepen our knowledge of molecular interactions and design therapeutics that are more effective. In this article, we have discussed the vicissitudes that the scientists have gone through to improve prediction accuracy, and examined the effective policies in predicting from different aspects, including the construction of high quality MSA, providing informative input features, and progresses in deep learning approaches. We have also briefly touched upon transitioning from predicting single-chain protein structures to predicting protein complex structures. Our findings point towards promoting open research environments to support the objectives of protein structure prediction.
Collapse
Affiliation(s)
- Faezeh Rahimzadeh
- Faculty of Electrical and Computer Engineering, University of Tabriz, Tabriz, Iran
| | | | - Pedram Salehpoor
- Faculty of Electrical and Computer Engineering, University of Tabriz, Tabriz, Iran
| | - Faegheh Golabi
- Department of Biomedical Engineering, Faculty of Advanced Medical Sciences, Tabriz University of Medical Sciences, Tabriz, Iran
| | - Shahin PourBahrami
- Department of Computer Engineering, Technical and Vocational University (TVU), Tehran, Iran
| |
Collapse
|
5
|
Gu X, Aranganathan A, Tiwary P. Empowering AlphaFold2 for protein conformation selective drug discovery with AlphaFold2-RAVE. ARXIV 2024:arXiv:2404.07102v3. [PMID: 38659642 PMCID: PMC11042445] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Subscribe] [Scholar Register] [Indexed: 04/26/2024]
Abstract
Small molecule drug design hinges on obtaining co-crystallized ligand-protein structures. Despite AlphaFold2's strides in protein native structure prediction, its focus on apo structures overlooks ligands and associated holo structures. Moreover, designing selective drugs often benefits from the targeting of diverse metastable conformations. Therefore, direct application of AlphaFold2 models in virtual screening and drug discovery remains tentative. Here, we demonstrate an AlphaFold2 based framework combined with all-atom enhanced sampling molecular dynamics and induced fit docking, named AF2RAVE-Glide, to conduct computational model based small molecule binding of metastable protein kinase conformations, initiated from protein sequences. We demonstrate the AF2RAVE-Glide workflow on three different protein kinases and their type I and II inhibitors, with special emphasis on binding of known type II kinase inhibitors which target the metastable classical DFG-out state. These states are not easy to sample from AlphaFold2. Here we demonstrate how with AF2RAVE these metastable conformations can be sampled for different kinases with high enough accuracy to enable subsequent docking of known type II kinase inhibitors with more than 50% success rates across docking calculations. We believe the protocol should be deployable for other kinases and more proteins generally.
Collapse
|
6
|
Wu F, Huang Y, Yang G, Ye S, Mukamel S, Jiang J. Unraveling dynamic protein structures by two-dimensional infrared spectra with a pretrained machine learning model. Proc Natl Acad Sci U S A 2024; 121:e2409257121. [PMID: 38917009 PMCID: PMC11228460 DOI: 10.1073/pnas.2409257121] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/09/2024] [Accepted: 05/28/2024] [Indexed: 06/27/2024] Open
Abstract
Dynamic protein structures are crucial for deciphering their diverse biological functions. Two-dimensional infrared (2DIR) spectroscopy stands as an ideal tool for tracing rapid conformational evolutions in proteins. However, linking spectral characteristics to dynamic structures poses a formidable challenge. Here, we present a pretrained machine learning model based on 2DIR spectra analysis. This model has learned signal features from approximately 204,300 spectra to establish a "spectrum-structure" correlation, thereby tracing the dynamic conformations of proteins. It excels in accurately predicting the dynamic content changes of various secondary structures and demonstrates universal transferability on real folding trajectories spanning timescales from microseconds to milliseconds. Beyond exceptional predictive performance, the model offers attention-based spectral explanations of dynamic conformational changes. Our 2DIR-based pretrained model is anticipated to provide unique insights into the dynamic structural information of proteins in their native environments.
Collapse
Affiliation(s)
- Fan Wu
- Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei230026, Anhui, China
| | - Yan Huang
- Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei230026, Anhui, China
| | - Guokun Yang
- Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei230026, Anhui, China
| | - Sheng Ye
- Anhui Provincial Engineering Research Center for Unmanned System and Intelligent Technology, School of Artificial Intelligence, Anhui University, Hefei230601, Anhui, China
| | - Shaul Mukamel
- Department of Chemistry and of Physics & Astronomy, University of California, Irvine, CA92697
| | - Jun Jiang
- Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei230026, Anhui, China
| |
Collapse
|
7
|
Basu S, Subedi U, Tonelli M, Afshinpour M, Tiwari N, Fuentes EJ, Chakravarty S. Assessing the functional roles of coevolving PHD finger residues. Protein Sci 2024; 33:e5065. [PMID: 38923615 PMCID: PMC11201814 DOI: 10.1002/pro.5065] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/31/2023] [Revised: 04/21/2024] [Accepted: 05/16/2024] [Indexed: 06/28/2024]
Abstract
Although in silico folding based on coevolving residue constraints in the deep-learning era has transformed protein structure prediction, the contributions of coevolving residues to protein folding, stability, and other functions in physical contexts remain to be clarified and experimentally validated. Herein, the PHD finger module, a well-known histone reader with distinct subtypes containing subtype-specific coevolving residues, was used as a model to experimentally assess the contributions of coevolving residues and to clarify their specific roles. The results of the assessment, including proteolysis and thermal unfolding of wildtype and mutant proteins, suggested that coevolving residues have varying contributions, despite their large in silico constraints. Residue positions with large constraints were found to contribute to stability in one subtype but not others. Computational sequence design and generative model-based energy estimates of individual structures were also implemented to complement the experimental assessment. Sequence design and energy estimates distinguish coevolving residues that contribute to folding from those that do not. The results of proteolytic analysis of mutations at positions contributing to folding were consistent with those suggested by sequence design and energy estimation. Thus, we report a comprehensive assessment of the contributions of coevolving residues, as well as a strategy based on a combination of approaches that should enable detailed understanding of the residue contributions in other large protein families.
Collapse
Affiliation(s)
- Shraddha Basu
- Department of Chemistry & BiochemistrySouth Dakota State UniversityBrookingsSouth DakotaUSA
| | - Ujwal Subedi
- Department of Chemistry & BiochemistrySouth Dakota State UniversityBrookingsSouth DakotaUSA
| | - Marco Tonelli
- National Magnetic Resonance Facility at Madison (NMRFAM), University of Wisconsin‐MadisonMadisonWisconsinUSA
| | - Maral Afshinpour
- Department of Chemistry & BiochemistrySouth Dakota State UniversityBrookingsSouth DakotaUSA
| | - Nitija Tiwari
- Department of Biochemistry & Molecular BiologyUniversity of IowaIowa CityIowaUSA
| | - Ernesto J. Fuentes
- Department of Biochemistry & Molecular BiologyUniversity of IowaIowa CityIowaUSA
| | - Suvobrata Chakravarty
- Department of Chemistry & BiochemistrySouth Dakota State UniversityBrookingsSouth DakotaUSA
| |
Collapse
|
8
|
Huang YJ, Montelione GT. Hidden Structural States of Proteins Revealed by Conformer Selection with AlphaFold-NMR. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.06.26.600902. [PMID: 38979209 PMCID: PMC11230435 DOI: 10.1101/2024.06.26.600902] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 07/10/2024]
Abstract
Recent advances in molecular modeling using deep learning can revolutionize our understanding of dynamic protein structures. NMR is particularly well-suited for determining dynamic features of biomolecular structures. The conventional process for determining biomolecular structures from experimental NMR data involves its representation as conformation-dependent restraints, followed by generation of structural models guided by these spatial restraints. Here we describe an alternative approach: generating a distribution of realistic protein conformational models using artificial intelligence-(AI-) based methods and then selecting the sets of conformers that best explain the experimental data. We applied this conformational selection approach to redetermine the solution NMR structure of the enzyme Gaussia luciferase. First, we generated a diverse set of conformer models using AlphaFold2 (AF2) with an enhanced sampling protocol. The models that best-fit NOESY and chemical shift data were then selected with a Bayesian scoring metric. The resulting models include features of both the published NMR structure and the standard AF2 model generated without enhanced sampling. This "AlphaFold-NMR" protocol also generated an alternative "open" conformational state that fits nearly as well to the overall NMR data but accounts for some NOESY data that is not consistent with first "closed" conformational state; while other NOESY data consistent with this second state are not consistent with the first conformational state. The structure of this "open" structural state differs from that of the "closed" state primarily by the position of a thumb-shaped loop between α-helices H5 and H6, revealing a cryptic surface pocket. These alternative conformational states of Gluc are supported by "double recall" analysis of NOESY data and AF2 models. Additional structural states are also indicated by backbone chemical shift data indicating partially-disordered conformations for the C-terminal segment. Considered as a multistate ensemble, these multiple states of Gluc together fit the NOESY and chemical shift data better than the "restraint-based" NMR structure and provide novel insights into its structure-dynamic-function relationships. This study demonstrates the potential of AI-based modeling with enhanced sampling to generate conformational ensembles followed by conformer selection with experimental data as an alternative to conventional restraint satisfaction protocols for protein NMR structure determination.
Collapse
Affiliation(s)
- Yuanpeng J Huang
- Dept of Chemistry and Chemical Biology, Center for Biotechnology and Interdisciplinary Sciences, Rensselaer Polytechnic Institute, Troy, New York, 12180 USA
| | - Gaetano T Montelione
- Dept of Chemistry and Chemical Biology, Center for Biotechnology and Interdisciplinary Sciences, Rensselaer Polytechnic Institute, Troy, New York, 12180 USA
| |
Collapse
|
9
|
Raisinghani N, Alshahrani M, Gupta G, Xiao S, Tao P, Verkhivker G. Exploring conformational landscapes and binding mechanisms of convergent evolution for the SARS-CoV-2 spike Omicron variant complexes with the ACE2 receptor using AlphaFold2-based structural ensembles and molecular dynamics simulations. Phys Chem Chem Phys 2024; 26:17720-17744. [PMID: 38869513 DOI: 10.1039/d4cp01372g] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 06/14/2024]
Abstract
In this study, we combined AlphaFold-based approaches for atomistic modeling of multiple protein states and microsecond molecular simulations to accurately characterize conformational ensembles evolution and binding mechanisms of convergent evolution for the SARS-CoV-2 spike Omicron variants BA.1, BA.2, BA.2.75, BA.3, BA.4/BA.5 and BQ.1.1. We employed and validated several different adaptations of the AlphaFold methodology for modeling of conformational ensembles including the introduced randomized full sequence scanning for manipulation of sequence variations to systematically explore conformational dynamics of Omicron spike protein complexes with the ACE2 receptor. Microsecond atomistic molecular dynamics (MD) simulations provide a detailed characterization of the conformational landscapes and thermodynamic stability of the Omicron variant complexes. By integrating the predictions of conformational ensembles from different AlphaFold adaptations and applying statistical confidence metrics we can expand characterization of the conformational ensembles and identify functional protein conformations that determine the equilibrium dynamics for the Omicron spike complexes with the ACE2. Conformational ensembles of the Omicron RBD-ACE2 complexes obtained using AlphaFold-based approaches for modeling protein states and MD simulations are employed for accurate comparative prediction of the binding energetics revealing an excellent agreement with the experimental data. In particular, the results demonstrated that AlphaFold-generated extended conformational ensembles can produce accurate binding energies for the Omicron RBD-ACE2 complexes. The results of this study suggested complementarities and potential synergies between AlphaFold predictions of protein conformational ensembles and MD simulations showing that integrating information from both methods can potentially yield a more adequate characterization of the conformational landscapes for the Omicron RBD-ACE2 complexes. This study provides insights in the interplay between conformational dynamics and binding, showing that evolution of Omicron variants through acquisition of convergent mutational sites may leverage conformational adaptability and dynamic couplings between key binding energy hotspots to optimize ACE2 binding affinity and enable immune evasion.
Collapse
Affiliation(s)
- Nishank Raisinghani
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, CA 92866, USA.
| | - Mohammed Alshahrani
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, CA 92866, USA.
| | - Grace Gupta
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, CA 92866, USA.
| | - Sian Xiao
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas, 75275, USA
| | - Peng Tao
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas, 75275, USA
| | - Gennady Verkhivker
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, CA 92866, USA.
- Department of Biomedical and Pharmaceutical Sciences, Chapman University School of Pharmacy, Irvine, CA 92618, USA
| |
Collapse
|
10
|
Raisinghani N, Alshahrani M, Gupta G, Tian H, Xiao S, Tao P, Verkhivker GM. Integration of a Randomized Sequence Scanning Approach in AlphaFold2 and Local Frustration Profiling of Conformational States Enable Interpretable Atomistic Characterization of Conformational Ensembles and Detection of Hidden Allosteric States in the ABL1 Protein Kinase. J Chem Theory Comput 2024; 20:5317-5336. [PMID: 38865109 DOI: 10.1021/acs.jctc.4c00222] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 06/13/2024]
Abstract
Despite the success of AlphaFold methods in predicting single protein structures, these methods showed intrinsic limitations in the characterization of multiple functional conformations of allosteric proteins. The recent NMR-based structural determination of the unbound ABL kinase in the active state and discovery of the inactive low-populated functional conformations that are unique for ABL kinase present an ideal challenge for the AlphaFold2 approaches. In the current study, we employ several adaptations of the AlphaFold2 methodology to predict protein conformational ensembles and allosteric states of the ABL kinase including randomized alanine sequence scanning combined with the multiple sequence alignment subsampling proposed in this study. We show that the proposed new AlphaFold2 adaptation combined with local frustration profiling of conformational states enables accurate prediction of the protein kinase structures and conformational ensembles, also offering a robust approach for interpretable characterization of the AlphaFold2 predictions and detection of hidden allosteric states. We found that the large high frustration residue clusters are uniquely characteristic of the low-populated, fully inactive ABL form and can define energetically frustrated cracking sites of conformational transitions, presenting difficult targets for AlphaFold2. The results of this study uncovered previously unappreciated fundamental connections between local frustration profiles of the functional allosteric states and the ability of AlphaFold2 methods to predict protein structural ensembles of the active and inactive states. This study showed that integration of the randomized sequence scanning adaptation of AlphaFold2 with a robust landscape-based analysis allows for interpretable atomistic predictions and characterization of protein conformational ensembles, providing a physical basis for the successes and limitations of current AlphaFold2 methods in detecting functional allosteric states that play a significant role in protein kinase regulation.
Collapse
Affiliation(s)
- Nishank Raisinghani
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
| | - Mohammed Alshahrani
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
| | - Grace Gupta
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
| | - Hao Tian
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas 75275, United States
| | - Sian Xiao
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas 75275, United States
| | - Peng Tao
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas 75275, United States
| | - Gennady M Verkhivker
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
- Department of Biomedical and Pharmaceutical Sciences, Chapman University School of Pharmacy, Irvine, California 92618, United States
- Department of Pharmacology, Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego, 9500 Gilman Drive, La Jolla, California 92093, United States
| |
Collapse
|
11
|
Duran C, Casadevall G, Osuna S. Harnessing conformational dynamics in enzyme catalysis to achieve nature-like catalytic efficiencies: the shortest path map tool for computational enzyme redesign. Faraday Discuss 2024. [PMID: 38910409 DOI: 10.1039/d3fd00156c] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 06/25/2024]
Abstract
Enzymes exhibit diverse conformations, as represented in the free energy landscape (FEL). Such conformational diversity provides enzymes with the ability to evolve towards novel functions. The challenge lies in identifying mutations that enhance specific conformational changes, especially if located in distal sites from the active site cavity. The shortest path map (SPM) method, which we developed to address this challenge, constructs a graph based on the distances and correlated motions of residues observed in nanosecond timescale molecular dynamics (MD) simulations. We recently introduced a template based AlphaFold2 (tAF2) approach coupled with 10 nanosecond MD simulations to quickly estimate the conformational landscape of enzymes and assess how the FEL is shifted after mutation. In this study, we evaluate the potential of SPM when coupled with tAF2-MD in estimating conformational heterogeneity and identifying key conformationally-relevant positions. The selected model system is the beta subunit of tryptophan synthase (TrpB). We compare how the SPM pathways differ when integrating tAF2 with different MD simulation lengths from as short as 10 ns until 50 ns and considering two distinct Amber forcefield and water models (ff14SB/TIP3P versus ff19SB/OPC). The new methodology can more effectively capture the distal mutations found in laboratory evolution, thus showcasing the efficacy of tAF2-MD-SPM in rapidly estimating enzyme dynamics and identifying the key conformationally relevant hotspots for computational enzyme engineering.
Collapse
Affiliation(s)
- Cristina Duran
- Departament de Química, Institut de Química Computacional i Catàlisi, Universitat de Girona, c/Maria Aurèlia Capmany 69, 17003, Girona, Spain.
| | - Guillem Casadevall
- Departament de Química, Institut de Química Computacional i Catàlisi, Universitat de Girona, c/Maria Aurèlia Capmany 69, 17003, Girona, Spain.
| | - Sílvia Osuna
- Departament de Química, Institut de Química Computacional i Catàlisi, Universitat de Girona, c/Maria Aurèlia Capmany 69, 17003, Girona, Spain.
- ICREA, Pg. Lluís Companys 23, 08010, Barcelona, Spain
| |
Collapse
|
12
|
Agarwal V, McShan AC. The power and pitfalls of AlphaFold2 for structure prediction beyond rigid globular proteins. Nat Chem Biol 2024:10.1038/s41589-024-01638-w. [PMID: 38907110 DOI: 10.1038/s41589-024-01638-w] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/22/2023] [Accepted: 04/29/2024] [Indexed: 06/23/2024]
Abstract
Artificial intelligence-driven advances in protein structure prediction in recent years have raised the question: has the protein structure-prediction problem been solved? Here, with a focus on nonglobular proteins, we highlight the many strengths and potential weaknesses of DeepMind's AlphaFold2 in the context of its biological and therapeutic applications. We summarize the subtleties associated with evaluation of AlphaFold2 model quality and reliability using the predicted local distance difference test (pLDDT) and predicted aligned error (PAE) values. We highlight various classes of proteins that AlphaFold2 can be applied to and the caveats involved. Concrete examples of how AlphaFold2 models can be integrated with experimental data in the form of small-angle X-ray scattering (SAXS), solution NMR, cryo-electron microscopy (cryo-EM) and X-ray diffraction are discussed. Finally, we highlight the need to move beyond structure prediction of rigid, static structural snapshots toward conformational ensembles and alternate biologically relevant states. The overarching theme is that careful consideration is due when using AlphaFold2-generated models to generate testable hypotheses and structural models, rather than treating predicted models as de facto ground truth structures.
Collapse
Affiliation(s)
- Vinayak Agarwal
- School of Chemistry and Biochemistry, Georgia Institute of Technology, Atlanta, GA, USA.
- School of Biological Sciences, Georgia Institute of Technology, Atlanta, GA, USA.
| | - Andrew C McShan
- School of Chemistry and Biochemistry, Georgia Institute of Technology, Atlanta, GA, USA.
| |
Collapse
|
13
|
Wayment-Steele HK, Otten R, Pitsawong W, Ojoawo AM, Glaser A, Calderone LA, Kern D. The conformational landscape of fold-switcher KaiB is tuned to the circadian rhythm timescale. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.06.03.597139. [PMID: 38895306 PMCID: PMC11185700 DOI: 10.1101/2024.06.03.597139] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 06/21/2024]
Abstract
How can a single protein domain encode a conformational landscape with multiple stably-folded states, and how do those states interconvert? Here, we use real-time and relaxation-dispersion NMR to characterize the conformational landscape of the circadian rhythm protein KaiB from Rhodobacter sphaeroides. Unique among known natural metamorphic proteins, this KaiB variant spontaneously interconverts between two monomeric states: the "Ground" and "Fold-switched" (FS) state. KaiB in its FS state interacts with multiple binding partners, including the central KaiC protein, to regulate circadian rhythms. We find that KaiB itself takes hours to interconvert between the Ground and FS state, underscoring the ability of a single sequence to encode the slow process needed for function. We reveal the rate-limiting step between the Ground and FS state is the cis-trans isomerization of three prolines in the fold-switching region by demonstrating interconversion acceleration by the prolyl isomerase CypA. The interconversion proceeds through a "partially disordered" (PD) state, where the C-terminal half becomes disordered while the N-terminal half remains stably folded. We discovered two additional properties of KaiB's landscape. Firstly, the Ground state experiences cold denaturation: at 4°C, the PD state becomes the majorly populated state. Secondly, the Ground state exchanges with a fourth state, the "Enigma" state, on the millisecond timescale. We combine AlphaFold2-based predictions and NMR chemical shift predictions to predict this "Enigma" state is a beta-strand register shift that eases buried charged residues, and support this structure experimentally. These results provide mechanistic insight in how evolution can design a single sequence that achieves specific timing needed for its function.
Collapse
Affiliation(s)
- Hannah K Wayment-Steele
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
| | - Renee Otten
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
- Present address: Treeline Biosciences, Watertown, MA, USA
| | - Warintra Pitsawong
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
- Present address: Biomolecular Discovery, Relay Therapeutics, Cambridge, MA, USA
| | - Adedolapo M Ojoawo
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
| | - Andrew Glaser
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
| | - Logan A Calderone
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
| | - Dorothee Kern
- Department of Biochemistry, Brandeis University and Howard Hughes Medical Institute, Waltham, MA, USA
| |
Collapse
|
14
|
Urvas L, Chiesa L, Bret G, Jacquemard C, Kellenberger E. Benchmarking AlphaFold-Generated Structures of Chemokine-Chemokine Receptor Complexes. J Chem Inf Model 2024; 64:4587-4600. [PMID: 38809680 DOI: 10.1021/acs.jcim.3c01835] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 05/31/2024]
Abstract
AlphaFold and AlphaFold-Multimer have become two essential tools for the modeling of unknown structures of proteins and protein complexes. In this work, we extensively benchmarked the quality of chemokine-chemokine receptor structures generated by AlphaFold-Multimer against experimentally determined structures. Our analysis considered both the global quality of the model, as well as key structural features for chemokine recognition. To study the effects of template and multiple sequence alignment parameters on the results, a new prediction pipeline called LIT-AlphaFold (https://github.com/LIT-CCM-lab/LIT-AlphaFold) was developed, allowing extensive input customization. AlphaFold-Multimer correctly predicted differences in chemokine binding orientation and accurately reproduced the unique binding orientation of the CXCL12-ACKR3 complex. Further, the predictions of the full receptor N-terminus provided insights into a putative chemokine recognition site 0.5. The accuracy of chemokine N-terminus binding mode prediction varied between complexes, but the confidence score permitted the distinguishing of residues that were very likely well positioned. Finally, we generated a high-confidence model of the unsolved CXCL12-CXCR4 complex, which agreed with experimental mutagenesis and cross-linking data.
Collapse
Affiliation(s)
- Lauri Urvas
- Laboratoire d'Innovation Thérapeutique, UMR 7200 CNRS, Université de Strasbourg, 67400 Illkirch, France
| | - Luca Chiesa
- Laboratoire d'Innovation Thérapeutique, UMR 7200 CNRS, Université de Strasbourg, 67400 Illkirch, France
| | - Guillaume Bret
- Laboratoire d'Innovation Thérapeutique, UMR 7200 CNRS, Université de Strasbourg, 67400 Illkirch, France
| | - Célien Jacquemard
- Laboratoire d'Innovation Thérapeutique, UMR 7200 CNRS, Université de Strasbourg, 67400 Illkirch, France
| | - Esther Kellenberger
- Laboratoire d'Innovation Thérapeutique, UMR 7200 CNRS, Université de Strasbourg, 67400 Illkirch, France
| |
Collapse
|
15
|
Giraldo R. The emergence of bacterial prions. PLoS Pathog 2024; 20:e1012253. [PMID: 38870093 PMCID: PMC11175392 DOI: 10.1371/journal.ppat.1012253] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 06/15/2024] Open
Affiliation(s)
- Rafael Giraldo
- Department of Microbial Biotechnology, National Center for Biotechnology (CNB-CSIC), Madrid, Spain
| |
Collapse
|
16
|
Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, Ronneberger O, Willmore L, Ballard AJ, Bambrick J, Bodenstein SW, Evans DA, Hung CC, O'Neill M, Reiman D, Tunyasuvunakool K, Wu Z, Žemgulytė A, Arvaniti E, Beattie C, Bertolli O, Bridgland A, Cherepanov A, Congreve M, Cowen-Rivers AI, Cowie A, Figurnov M, Fuchs FB, Gladman H, Jain R, Khan YA, Low CMR, Perlin K, Potapenko A, Savy P, Singh S, Stecula A, Thillaisundaram A, Tong C, Yakneen S, Zhong ED, Zielinski M, Žídek A, Bapst V, Kohli P, Jaderberg M, Hassabis D, Jumper JM. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024; 630:493-500. [PMID: 38718835 PMCID: PMC11168924 DOI: 10.1038/s41586-024-07487-w] [Citation(s) in RCA: 10] [Impact Index Per Article: 10.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/19/2023] [Accepted: 04/29/2024] [Indexed: 06/13/2024]
Abstract
The introduction of AlphaFold 21 has spurred a revolution in modelling the structure of proteins and their interactions, enabling a huge range of applications in protein modelling and design2-6. Here we describe our AlphaFold 3 model with a substantially updated diffusion-based architecture that is capable of predicting the joint structure of complexes including proteins, nucleic acids, small molecules, ions and modified residues. The new AlphaFold model demonstrates substantially improved accuracy over many previous specialized tools: far greater accuracy for protein-ligand interactions compared with state-of-the-art docking tools, much higher accuracy for protein-nucleic acid interactions compared with nucleic-acid-specific predictors and substantially higher antibody-antigen prediction accuracy compared with AlphaFold-Multimer v.2.37,8. Together, these results show that high-accuracy modelling across biomolecular space is possible within a single unified deep-learning framework.
Collapse
Affiliation(s)
| | - Jonas Adler
- Core Contributor, Google DeepMind, London, UK
| | - Jack Dunger
- Core Contributor, Google DeepMind, London, UK
| | | | - Tim Green
- Core Contributor, Google DeepMind, London, UK
| | | | | | | | | | | | | | | | | | | | | | | | - Zachary Wu
- Core Contributor, Google DeepMind, London, UK
| | | | | | | | | | | | | | | | | | | | | | | | | | | | - Yousuf A Khan
- Google DeepMind, London, UK
- Department of Molecular and Cellular Physiology, Stanford University, Stanford, CA, USA
| | | | | | | | | | | | | | | | | | | | - Ellen D Zhong
- Google DeepMind, London, UK
- Department of Computer Science, Princeton University, Princeton, NJ, USA
| | | | | | | | | | | | - Demis Hassabis
- Core Contributor, Google DeepMind, London, UK.
- Core Contributor, Isomorphic Labs, London, UK.
| | | |
Collapse
|
17
|
Porter LL, Artsimovitch I, Ramírez-Sarmiento CA. Metamorphic proteins and how to find them. Curr Opin Struct Biol 2024; 86:102807. [PMID: 38537533 PMCID: PMC11102287 DOI: 10.1016/j.sbi.2024.102807] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 02/01/2024] [Revised: 03/05/2024] [Accepted: 03/06/2024] [Indexed: 04/04/2024]
Abstract
In the last two decades, our existing notion that most foldable proteins have a unique native state has been challenged by the discovery of metamorphic proteins, which reversibly interconvert between multiple, sometimes highly dissimilar, native states. As the number of known metamorphic proteins increases, several computational and experimental strategies have emerged for gaining insights about their refolding processes and identifying unknown metamorphic proteins amongst the known proteome. In this review, we describe the current advances in biophysically and functionally ascertaining the structural interconversions of metamorphic proteins and how coevolution can be harnessed to identify novel metamorphic proteins from sequence information. We also discuss the challenges and ongoing efforts in using artificial intelligence-based protein structure prediction methods to discover metamorphic proteins and predict their corresponding three-dimensional structures.
Collapse
Affiliation(s)
- Lauren L Porter
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA; Biochemistry and Biophysics Center, National Heart, Lung, and Blood Institute, National Institutes of Health, Bethesda, MD 20892, USA.
| | - Irina Artsimovitch
- Department of Microbiology and Center for RNA Biology, The Ohio State University, Columbus, OH 43210, USA.
| | - César A Ramírez-Sarmiento
- Institute for Biological and Medical Engineering, Schools of Engineering, Medicine and Biological Sciences, Pontificia Universidad Católica de Chile, Santiago 7820436, Chile; ANID, Millennium Science Initiative Program, Millennium Institute for Integrative Biology (iBio), Santiago 833150, Chile.
| |
Collapse
|
18
|
Shor B, Schneidman-Duhovny D. Integrative modeling meets deep learning: Recent advances in modeling protein assemblies. Curr Opin Struct Biol 2024; 87:102841. [PMID: 38795564 DOI: 10.1016/j.sbi.2024.102841] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 02/29/2024] [Revised: 04/24/2024] [Accepted: 04/27/2024] [Indexed: 05/28/2024]
Abstract
Recent progress in protein structure prediction based on deep learning revolutionized the field of Structural Biology. Beyond single proteins, it also enabled high-throughput prediction of structures of protein-protein interactions. Despite the success in predicting complex structures, large macromolecular assemblies still require specialized approaches. Here we describe recent advances in modeling macromolecular assemblies using integrative and hierarchical approaches. We highlight applications that predict protein-protein interactions and challenges in modeling complexes based on the interaction networks, including the prediction of complex stoichiometry and heterogeneity.
Collapse
Affiliation(s)
- Ben Shor
- The Rachel and Selim Benin School of Computer Science and Engineering, The Hebrew University of Jerusalem, Jerusalem, Israel. https://twitter.com/ben_shor
| | - Dina Schneidman-Duhovny
- The Rachel and Selim Benin School of Computer Science and Engineering, The Hebrew University of Jerusalem, Jerusalem, Israel.
| |
Collapse
|
19
|
Raisinghani N, Alshahrani M, Gupta G, Tian H, Xiao S, Tao P, Verkhivker G. Prediction of Conformational Ensembles and Structural Effects of State-Switching Allosteric Mutants in the Protein Kinases Using Comparative Analysis of AlphaFold2 Adaptations with Sequence Masking and Shallow Subsampling. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.05.17.594786. [PMID: 38798650 PMCID: PMC11118581 DOI: 10.1101/2024.05.17.594786] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 05/29/2024]
Abstract
Despite the success of AlphaFold2 approaches in predicting single protein structures, these methods showed intrinsic limitations in predicting multiple functional conformations of allosteric proteins and have been challenged to accurately capture of the effects of single point mutations that induced significant structural changes. We systematically examined several implementations of AlphaFold2 methods to predict conformational ensembles for state-switching mutants of the ABL kinase. The results revealed that a combination of randomized alanine sequence masking with shallow multiple sequence alignment subsampling can significantly expand the conformational diversity of the predicted structural ensembles and capture shifts in populations of the active and inactive ABL states. Consistent with the NMR experiments, the predicted conformational ensembles for M309L/L320I and M309L/H415P ABL mutants that perturb the regulatory spine networks featured the increased population of the fully closed inactive state. On the other hand, the predicted conformational ensembles for the G269E/M309L/T334I and M309L/L320I/T334I triple ABL mutants that share activating T334I gate-keeper substitution are dominated by the active ABL form. The proposed adaptation of AlphaFold can reproduce the experimentally observed mutation-induced redistributions in the relative populations of the active and inactive ABL states and capture the effects of regulatory mutations on allosteric structural rearrangements of the kinase domain. The ensemble-based network analysis complemented AlphaFold predictions by revealing allosteric mediating centers that often directly correspond to state-switching mutational sites or reside in their immediate local structural proximity, which may explain the global effect of regulatory mutations on structural changes between the ABL states. This study suggested that attention-based learning of long-range dependencies between sequence positions in homologous folds and deciphering patterns of allosteric interactions may further augment the predictive abilities of AlphaFold methods for modeling of alternative protein sates, conformational ensembles and mutation-induced structural transformations.
Collapse
|
20
|
Raisinghani N, Alshahrani M, Gupta G, Xiao S, Tao P, Verkhivker G. AlphaFold2 Predictions of Conformational Ensembles and Atomistic Simulations of the SARS-CoV-2 Spike XBB Lineages Reveal Epistatic Couplings between Convergent Mutational Hotspots that Control ACE2 Affinity. J Phys Chem B 2024; 128:4696-4715. [PMID: 38696745 DOI: 10.1021/acs.jpcb.4c01341] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 05/04/2024]
Abstract
In this study, we combined AlphaFold-based atomistic structural modeling, microsecond molecular simulations, mutational profiling, and network analysis to characterize binding mechanisms of the SARS-CoV-2 spike protein with the host receptor ACE2 for a series of Omicron XBB variants including XBB.1.5, XBB.1.5+L455F, XBB.1.5+F456L, and XBB.1.5+L455F+F456L. AlphaFold-based structural and dynamic modeling of SARS-CoV-2 Spike XBB lineages can accurately predict the experimental structures and characterize conformational ensembles of the spike protein complexes with the ACE2. Microsecond molecular dynamics simulations identified important differences in the conformational landscapes and equilibrium ensembles of the XBB variants, suggesting that combining AlphaFold predictions of multiple conformations with molecular dynamics simulations can provide a complementary approach for the characterization of functional protein states and binding mechanisms. Using the ensemble-based mutational profiling of protein residues and physics-based rigorous calculations of binding affinities, we identified binding energy hotspots and characterized the molecular basis underlying epistatic couplings between convergent mutational hotspots. Consistent with the experiments, the results revealed the mediating role of the Q493 hotspot in the synchronization of epistatic couplings between L455F and F456L mutations, providing a quantitative insight into the energetic determinants underlying binding differences between XBB lineages. We also proposed a network-based perturbation approach for mutational profiling of allosteric communications and uncovered the important relationships between allosteric centers mediating long-range communication and binding hotspots of epistatic couplings. The results of this study support a mechanism in which the binding mechanisms of the XBB variants may be determined by epistatic effects between convergent evolutionary hotspots that control ACE2 binding.
Collapse
Affiliation(s)
- Nishank Raisinghani
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
| | - Mohammed Alshahrani
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
| | - Grace Gupta
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
| | - Sian Xiao
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas 75275, United States
| | - Peng Tao
- Department of Chemistry, Center for Research Computing, Center for Drug Discovery, Design, and Delivery (CD4), Southern Methodist University, Dallas, Texas 75275, United States
| | - Gennady Verkhivker
- Keck Center for Science and Engineering, Graduate Program in Computational and Data Sciences, Schmid College of Science and Technology, Chapman University, Orange, California 92866, United States
- Department of Biomedical and Pharmaceutical Sciences, Chapman University School of Pharmacy, Irvine, California 92618, United States
| |
Collapse
|
21
|
Ahdritz G, Bouatta N, Floristean C, Kadyan S, Xia Q, Gerecke W, O'Donnell TJ, Berenberg D, Fisk I, Zanichelli N, Zhang B, Nowaczynski A, Wang B, Stepniewska-Dziubinska MM, Zhang S, Ojewole A, Guney ME, Biderman S, Watkins AM, Ra S, Lorenzo PR, Nivon L, Weitzner B, Ban YEA, Chen S, Zhang M, Li C, Song SL, He Y, Sorger PK, Mostaque E, Zhang Z, Bonneau R, AlQuraishi M. OpenFold: retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization. Nat Methods 2024:10.1038/s41592-024-02272-z. [PMID: 38744917 DOI: 10.1038/s41592-024-02272-z] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/14/2023] [Accepted: 04/03/2024] [Indexed: 05/16/2024]
Abstract
AlphaFold2 revolutionized structural biology with the ability to predict protein structures with exceptionally high accuracy. Its implementation, however, lacks the code and data required to train new models. These are necessary to (1) tackle new tasks, like protein-ligand complex structure prediction, (2) investigate the process by which the model learns and (3) assess the model's capacity to generalize to unseen regions of fold space. Here we report OpenFold, a fast, memory efficient and trainable implementation of AlphaFold2. We train OpenFold from scratch, matching the accuracy of AlphaFold2. Having established parity, we find that OpenFold is remarkably robust at generalizing even when the size and diversity of its training set is deliberately limited, including near-complete elisions of classes of secondary structure elements. By analyzing intermediate structures produced during training, we also gain insights into the hierarchical manner in which OpenFold learns to fold. In sum, our studies demonstrate the power and utility of OpenFold, which we believe will prove to be a crucial resource for the protein modeling community.
Collapse
Affiliation(s)
- Gustaf Ahdritz
- Department of Systems Biology, Columbia University, New York, NY, USA
- Harvard University, Cambridge, MA, USA
| | - Nazim Bouatta
- Laboratory of Systems Pharmacology, Harvard Medical School, Boston, MA, USA.
| | | | - Sachin Kadyan
- Department of Systems Biology, Columbia University, New York, NY, USA
| | - Qinghui Xia
- Department of Systems Biology, Columbia University, New York, NY, USA
| | - William Gerecke
- Laboratory of Systems Pharmacology, Harvard Medical School, Boston, MA, USA
| | | | - Daniel Berenberg
- Department of Computer Science, Courant Institute of Mathematical Sciences, New York University, New York, NY, USA
| | - Ian Fisk
- Flatiron Institute, New York, NY, USA
| | | | - Bo Zhang
- Scientific Computing and Imaging Institute, University of Utah, Salt Lake City, UT, USA
| | | | | | | | | | | | | | - Stella Biderman
- EleutherAI, New York, NY, USA
- Booz Allen Hamilton, McLean, VA, USA
| | | | - Stephen Ra
- Prescient Design, Genentech, New York, NY, USA
| | | | | | | | | | | | - Minjia Zhang
- University of Illinois at Urbana-Champaign, Champaign, IL, USA
| | | | | | | | - Peter K Sorger
- Laboratory of Systems Pharmacology, Harvard Medical School, Boston, MA, USA
| | | | - Zhao Zhang
- Rutgers University, New Brunswick, NJ, USA
| | | | | |
Collapse
|
22
|
Guo HB, Huntington B, Perminov A, Smith K, Hastings N, Dennis P, Kelley-Loughnane N, Berry R. AlphaFold2 modeling and molecular dynamics simulations of an intrinsically disordered protein. PLoS One 2024; 19:e0301866. [PMID: 38739602 PMCID: PMC11090348 DOI: 10.1371/journal.pone.0301866] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/05/2023] [Accepted: 03/23/2024] [Indexed: 05/16/2024] Open
Abstract
We use AlphaFold2 (AF2) to model the monomer and dimer structures of an intrinsically disordered protein (IDP), Nvjp-1, assisted by molecular dynamics (MD) simulations. We observe relatively rigid dimeric structures of Nvjp-1 when compared with the monomer structures. We suggest that protein conformations from multiple AF2 models and those from MD trajectories exhibit a coherent trend: the conformations of an IDP are deviated from each other and the conformations of a well-folded protein are consistent with each other. We use a residue-residue interaction network (RIN) derived from the contact map which show that the residue-residue interactions in Nvjp-1 are mainly transient; however, those in a well-folded protein are mainly persistent. Despite the variation in 3D shapes, we show that the AF2 models of both disordered and ordered proteins exhibit highly consistent profiles of the pLDDT (predicted local distance difference test) scores. These results indicate a potential protocol to justify the IDPs based on multiple AF2 models and MD simulations.
Collapse
Affiliation(s)
- Hao-Bo Guo
- Material and Manufacturing Directorate, Air Force Research Laboratory, WPAFB, Mason, OH, United States of America
- UES Inc., Dayton, OH, United States of America
| | - Baxter Huntington
- Material and Manufacturing Directorate, Air Force Research Laboratory, WPAFB, Mason, OH, United States of America
- Miami University, Oxford, OH, United States of America
| | - Alexander Perminov
- Material and Manufacturing Directorate, Air Force Research Laboratory, WPAFB, Mason, OH, United States of America
- Miami University, Oxford, OH, United States of America
| | - Kenya Smith
- United States Air Force Academy, Colorado Springs, CO, United States of America
| | - Nicholas Hastings
- United States Air Force Academy, Colorado Springs, CO, United States of America
| | - Patrick Dennis
- Material and Manufacturing Directorate, Air Force Research Laboratory, WPAFB, Mason, OH, United States of America
| | - Nancy Kelley-Loughnane
- Material and Manufacturing Directorate, Air Force Research Laboratory, WPAFB, Mason, OH, United States of America
| | - Rajiv Berry
- Material and Manufacturing Directorate, Air Force Research Laboratory, WPAFB, Mason, OH, United States of America
| |
Collapse
|
23
|
Zhang O, Naik SA, Liu ZH, Forman-Kay J, Head-Gordon T. A Curated Rotamer Library for Common Post-Translational Modifications of Proteins. ARXIV 2024:arXiv:2405.03120v1. [PMID: 38764597 PMCID: PMC11100909] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Grants] [Subscribe] [Scholar Register] [Indexed: 05/21/2024]
Abstract
Sidechain rotamer libraries of the common amino acids of a protein are useful for folded protein structure determination and for generating ensembles of intrinsically disordered proteins (IDPs). However much of protein function is modulated beyond the translated sequence through thFiguree introduction of post-translational modifications (PTMs). In this work we have provided a curated set of side chain rotamers for the most common PTMs derived from the RCSB PDB database, including phosphorylated, methylated, and acetylated sidechains. Our rotamer libraries improve upon existing methods such as SIDEpro and Rosetta in predicting the experimental structures for PTMs in folded proteins. In addition, we showcase our PTM libraries in full use by generating ensembles with the Monte Carlo Side Chain Entropy (MCSCE) for folded proteins, and combining MCSCE with the Local Disordered Region Sampling algorithms within IDPConformerGenerator for proteins with intrinsically disordered regions.
Collapse
Affiliation(s)
- Oufan Zhang
- Kenneth S. Pitzer Center for Theoretical Chemistry, University of California, Berkeley, Berkeley, California 94720, USA
| | - Shubhankar A. Naik
- Department of Chemistry, University of California, Berkeley, Berkeley, California 94720, USA
| | - Zi Hao Liu
- Molecular Medicine Program, Hospital for Sick Children, Toronto, Ontario M5G 0A4, Canada
- Department of Biochemistry, University of Toronto, Toronto, Ontario M5S 1A8, Canada
| | - Julie Forman-Kay
- Molecular Medicine Program, Hospital for Sick Children, Toronto, Ontario M5G 0A4, Canada
- Department of Biochemistry, University of Toronto, Toronto, Ontario M5S 1A8, Canada
| | - Teresa Head-Gordon
- Kenneth S. Pitzer Center for Theoretical Chemistry, University of California, Berkeley, Berkeley, California 94720, USA
- Department of Chemistry, University of California, Berkeley, Berkeley, California 94720, USA
- Department of Bioengineering, University of California, Berkeley, Berkeley, California 94720, USA
- Department of Chemical and Biomolecular Engineering, University of California, Berkeley, Berkeley, California 94720, USA
| |
Collapse
|
24
|
Renthal R. Arthropod repellent interactions with olfactory receptors and ionotropic receptors analyzed by molecular modeling. CURRENT RESEARCH IN INSECT SCIENCE 2024; 5:100082. [PMID: 38765913 PMCID: PMC11101704 DOI: 10.1016/j.cris.2024.100082] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Subscribe] [Scholar Register] [Received: 12/23/2023] [Revised: 04/29/2024] [Accepted: 05/01/2024] [Indexed: 05/22/2024]
Abstract
The main insect chemoreceptors are olfactory receptors (ORs), gustatory receptors (GRs) and ionotropic receptors (IRs). The odorant binding sites of many insect ORs appear to be occluded and inaccessible from the surface of the receptor protein, based on the three-dimensional structure of OR5 from the jumping bristletail Machilis hrabei (MhraOR5) and a survey of a sample of vinegar fly (Drosophila melanogaster) OR structures obtained from artificial intellegence (A.I.) modeling. Molecular dynamics simulations revealed that the occluded site can become accessible through tunnels that transiently open and close. The present study extends this analysis to examine seventeen ORs and one GR docking with ligands that have known valence: nine that signal attraction and nine that signal aversion. All but one of the receptors displayed occluded ligand binding sites analogous to MhraOR5, and docking software predicted the known attractant and repellent ligands will bind to the occluded sites. Docking of the repellent DEET was examined, and more than half of the OR ligand sites were predicted to bind DEET, including receptors that signal aversion as well as those that signal attraction. However, DEET may not actually have access to all the attractant binding sites. The larger size and lower flexibility of repellent molecules may restrict their passage through the tunnel bottlenecks, which could act as filters to select access to the ligand binding sites. In contrast to ORs and GRs, the IR ligand binding site is in an extracellular domain known to undergo a large conformational change from an open to a closed state. A.I. models of two D. melanogaster IRs of known valence and two blacklegged tick (Ixodes scapularis) IRs having unknown ligands were computationally tested for attractant and repellent binding. The ligand-binding sites in the closed state appear inaccessible to the protein surface, so attractants and repellents must bind initially at an accessible site in the open state before triggering the conformational change. In some IRs, repellent binding sites were identified at exterior sites adjacent to the ligand-binding site. These may be allosteric sites that, when occupied by repellents, can stabilize the open state of an attractant IR, or stabilize the closed state of an IR in the absence of its activating ligand. The model of D. melanogaster IR64a suggests a possible molecular mechanism for the activation of this IR by H+. The amino acids involved in this proposed mechanism are conserved in IR64a from several Dipteran pest species and disease vectors, potentially offering a route to discovery of new repellents that act via the allosteric site.
Collapse
Affiliation(s)
- Robert Renthal
- Department of Neuroscience, Developmental and Regenerative Biology, University of Texas at San Antonio, San Antonio, TX, United States
| |
Collapse
|
25
|
Xie P, Xu Y, Tang J, Wu S, Gao H. Multifaceted regulation of siderophore synthesis by multiple regulatory systems in Shewanella oneidensis. Commun Biol 2024; 7:498. [PMID: 38664541 PMCID: PMC11045786 DOI: 10.1038/s42003-024-06193-7] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/15/2024] [Accepted: 04/15/2024] [Indexed: 04/28/2024] Open
Abstract
Siderophore-dependent iron uptake is a mechanism by which microorganisms scavenge and utilize iron for their survival, growth, and many specialized activities, such as pathogenicity. The siderophore biosynthetic system PubABC in Shewanella can synthesize a series of distinct siderophores, yet how it is regulated in response to iron availability remains largely unexplored. Here, by whole genome screening we identify TCS components histidine kinase (HK) BarA and response regulator (RR) SsoR as positive regulators of siderophore biosynthesis. While BarA partners with UvrY to mediate expression of pubABC post-transcriptionally via the Csr regulatory cascade, SsoR is an atypical orphan RR of the OmpR/PhoB subfamily that activates transcription in a phosphorylation-independent manner. By combining structural analysis and molecular dynamics simulations, we observe conformational changes in OmpR/PhoB-like RRs that illustrate the impact of phosphorylation on dynamic properties, and that SsoR is locked in the 'phosphorylated' state found in phosphorylation-dependent counterparts of the same subfamily. Furthermore, we show that iron homeostasis global regulator Fur, in addition to mediating transcription of its own regulon, acts as the sensor of iron starvation to increase SsoR production when needed. Overall, this study delineates an intricate, multi-tiered transcriptional and post-transcriptional regulatory network that governs siderophore biosynthesis.
Collapse
Affiliation(s)
- Peilu Xie
- Institute of Microbiology and College of Life Sciences, Zhejiang University, Hangzhou, Zhejiang, 310058, China
| | - Yuanyou Xu
- Institute of Microbiology and College of Life Sciences, Zhejiang University, Hangzhou, Zhejiang, 310058, China
| | - Jiaxin Tang
- Institute of Microbiology and College of Life Sciences, Zhejiang University, Hangzhou, Zhejiang, 310058, China
| | - Shihua Wu
- Institute of Microbiology and College of Life Sciences, Zhejiang University, Hangzhou, Zhejiang, 310058, China.
| | - Haichun Gao
- Institute of Microbiology and College of Life Sciences, Zhejiang University, Hangzhou, Zhejiang, 310058, China.
| |
Collapse
|
26
|
Xie T, Huang J. Can Protein Structure Prediction Methods Capture Alternative Conformations of Membrane Transporters? J Chem Inf Model 2024; 64:3524-3536. [PMID: 38564295 DOI: 10.1021/acs.jcim.3c01936] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 04/04/2024]
Abstract
Understanding the conformational dynamics of proteins, such as the inward-facing (IF) and outward-facing (OF) transition observed in transporters, is vital for elucidating their functional mechanisms. Despite significant advances in protein structure prediction (PSP) over the past three decades, most efforts have been focused on single-state prediction, leaving multistate or alternative conformation prediction (ACP) relatively unexplored. This discrepancy has led to the development of highly accurate PSP methods such as AlphaFold, yet their capabilities for ACP remain limited. To investigate the performance of current PSP methods in ACP, we curated a data set, named IOMemP, consisting of 32 experimentally determined high-resolution IF and OF structures of 16 membrane proteins with substantial conformational changes. We benchmarked 12 representative PSP methods, along with two recent multistate methods based on AlphaFold, against this data set. Our findings reveal a remarkably consistent preference for specific states across various PSP methods. We elucidated how coevolution information in MSAs influences state preference. Moreover, we showed that AlphaFold, when excluding coevolution information, estimated similar energies between the experimental IF and OF conformations, indicating that the energy model learned by AlphaFold is not biased toward any particular state. Our IOMemP data set and benchmark results are anticipated to advance the development of robust ACP methods.
Collapse
Affiliation(s)
- Tengyu Xie
- College of Life Science, Zhejiang University, HangZhou Zhejiang 310058, China
- Key Laboratory of Structural Biology of Zhejiang Province, School of Life Sciences, Westlake University, HangZhou Zhejiang 310024, China
- Westlake AI Therapeutics Lab, Westlake Laboratory of Life Sciences and Biomedicine, HangZhou Zhejiang 310024, China
| | - Jing Huang
- College of Life Science, Zhejiang University, HangZhou Zhejiang 310058, China
- Key Laboratory of Structural Biology of Zhejiang Province, School of Life Sciences, Westlake University, HangZhou Zhejiang 310024, China
- Westlake AI Therapeutics Lab, Westlake Laboratory of Life Sciences and Biomedicine, HangZhou Zhejiang 310024, China
| |
Collapse
|
27
|
Raisinghani N, Alshahrani M, Gupta G, Xiao S, Tao P, Verkhivker G. Predicting Functional Conformational Ensembles and Binding Mechanisms of Convergent Evolution for SARS-CoV-2 Spike Omicron Variants Using AlphaFold2 Sequence Scanning Adaptations and Molecular Dynamics Simulations. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.04.02.587850. [PMID: 38617283 PMCID: PMC11014522 DOI: 10.1101/2024.04.02.587850] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 04/16/2024]
Abstract
In this study, we combined AlphaFold-based approaches for atomistic modeling of multiple protein states and microsecond molecular simulations to accurately characterize conformational ensembles and binding mechanisms of convergent evolution for the SARS-CoV-2 Spike Omicron variants BA.1, BA.2, BA.2.75, BA.3, BA.4/BA.5 and BQ.1.1. We employed and validated several different adaptations of the AlphaFold methodology for modeling of conformational ensembles including the introduced randomized full sequence scanning for manipulation of sequence variations to systematically explore conformational dynamics of Omicron Spike protein complexes with the ACE2 receptor. Microsecond atomistic molecular dynamic simulations provide a detailed characterization of the conformational landscapes and thermodynamic stability of the Omicron variant complexes. By integrating the predictions of conformational ensembles from different AlphaFold adaptations and applying statistical confidence metrics we can expand characterization of the conformational ensembles and identify functional protein conformations that determine the equilibrium dynamics for the Omicron Spike complexes with the ACE2. Conformational ensembles of the Omicron RBD-ACE2 complexes obtained using AlphaFold-based approaches for modeling protein states and molecular dynamics simulations are employed for accurate comparative prediction of the binding energetics revealing an excellent agreement with the experimental data. In particular, the results demonstrated that AlphaFold-generated extended conformational ensembles can produce accurate binding energies for the Omicron RBD-ACE2 complexes. The results of this study suggested complementarities and potential synergies between AlphaFold predictions of protein conformational ensembles and molecular dynamics simulations showing that integrating information from both methods can potentially yield a more adequate characterization of the conformational landscapes for the Omicron RBD-ACE2 complexes. This study provides insights in the interplay between conformational dynamics and binding, showing that evolution of Omicron variants through acquisition of convergent mutational sites may leverage conformational adaptability and dynamic couplings between key binding energy hotspots to optimize ACE2 binding affinity and enable immune evasion.
Collapse
|
28
|
Monté D, Lens Z, Dewitte F, Villeret V, Verger A. Assessment of machine-learning predictions for the Mediator complex subunit MED25 ACID domain interactions with transactivation domains. FEBS Lett 2024; 598:758-773. [PMID: 38436147 DOI: 10.1002/1873-3468.14837] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 11/30/2023] [Revised: 02/01/2024] [Accepted: 02/10/2024] [Indexed: 03/05/2024]
Abstract
The human Mediator complex subunit MED25 binds transactivation domains (TADs) present in various cellular and viral proteins using two binding interfaces, named H1 and H2, which are found on opposite sides of its ACID domain. Here, we use and compare deep learning methods to characterize human MED25-TAD interfaces and assess the predicted models to published experimental data. For the H1 interface, AlphaFold produces predictions with high-reliability scores that agree well with experimental data, while the H2 interface predictions appear inconsistent, preventing reliable binding modes. Despite these limitations, we experimentally assess the validity of MED25 interface predictions with the viral transcriptional activators Lana-1 and IE62. AlphaFold predictions also suggest the existence of a unique hydrophobic pocket for the Arabidopsis MED25 ACID domain.
Collapse
Affiliation(s)
- Didier Monté
- CNRS EMR 9002 Integrative Structural Biology, Inserm U 1167 - RID-AGE, Univ. Lille, CHU Lille, Institut Pasteur de Lille, France
| | - Zoé Lens
- CNRS EMR 9002 Integrative Structural Biology, Inserm U 1167 - RID-AGE, Univ. Lille, CHU Lille, Institut Pasteur de Lille, France
| | - Frédérique Dewitte
- CNRS EMR 9002 Integrative Structural Biology, Inserm U 1167 - RID-AGE, Univ. Lille, CHU Lille, Institut Pasteur de Lille, France
| | - Vincent Villeret
- CNRS EMR 9002 Integrative Structural Biology, Inserm U 1167 - RID-AGE, Univ. Lille, CHU Lille, Institut Pasteur de Lille, France
| | - Alexis Verger
- CNRS EMR 9002 Integrative Structural Biology, Inserm U 1167 - RID-AGE, Univ. Lille, CHU Lille, Institut Pasteur de Lille, France
| |
Collapse
|
29
|
Monteiro da Silva G, Cui JY, Dalgarno DC, Lisi GP, Rubenstein BM. High-throughput prediction of protein conformational distributions with subsampled AlphaFold2. Nat Commun 2024; 15:2464. [PMID: 38538622 PMCID: PMC10973385 DOI: 10.1038/s41467-024-46715-9] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/03/2023] [Accepted: 02/28/2024] [Indexed: 04/12/2024] Open
Abstract
This paper presents an innovative approach for predicting the relative populations of protein conformations using AlphaFold 2, an AI-powered method that has revolutionized biology by enabling the accurate prediction of protein structures. While AlphaFold 2 has shown exceptional accuracy and speed, it is designed to predict proteins' ground state conformations and is limited in its ability to predict conformational landscapes. Here, we demonstrate how AlphaFold 2 can directly predict the relative populations of different protein conformations by subsampling multiple sequence alignments. We tested our method against nuclear magnetic resonance experiments on two proteins with drastically different amounts of available sequence data, Abl1 kinase and the granulocyte-macrophage colony-stimulating factor, and predicted changes in their relative state populations with more than 80% accuracy. Our subsampling approach worked best when used to qualitatively predict the effects of mutations or evolution on the conformational landscape and well-populated states of proteins. It thus offers a fast and cost-effective way to predict the relative populations of protein conformations at even single-point mutation resolution, making it a useful tool for pharmacology, analysis of experimental results, and predicting evolution.
Collapse
Affiliation(s)
| | - Jennifer Y Cui
- Brown University Department of Molecular and Cell Biology and Biochemistry, Providence, RI, USA
| | | | - George P Lisi
- Brown University Department of Molecular and Cell Biology and Biochemistry, Providence, RI, USA
- Brown University Department of Chemistry, Providence, RI, USA
| | - Brenda M Rubenstein
- Brown University Department of Molecular and Cell Biology and Biochemistry, Providence, RI, USA.
- Brown University Department of Chemistry, Providence, RI, USA.
| |
Collapse
|
30
|
Yang S, Song C. Switching Go̅ -Martini for Investigating Protein Conformational Transitions and Associated Protein-Lipid Interactions. J Chem Theory Comput 2024; 20:2618-2629. [PMID: 38447049 DOI: 10.1021/acs.jctc.3c01222] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 03/08/2024]
Abstract
Proteins are dynamic biomolecules that can transform between different conformational states when exerting physiological functions, which is difficult to simulate using all-atom methods. Coarse-grained (CG) Go̅-like models are widely used to investigate large-scale conformational transitions, which usually adopt implicit solvent models and therefore cannot explicitly capture the interaction between proteins and surrounding molecules, such as water and lipid molecules. Here, we present a new method, named Switching Go̅-Martini, to simulate large-scale protein conformational transitions between different states, based on the switching Go̅ method and the CG Martini 3 force field. The method is straightforward and efficient, as demonstrated by the benchmarking applications for multiple protein systems, including glutamine binding protein (GlnBP), adenylate kinase (AdK), and β2-adrenergic receptor (β2AR). Moreover, by employing the Switching Go̅-Martini method, we can not only unveil the conformational transition from the E2Pi-PL state to E1 state of the type 4 P-type ATPase (P4-ATPase) flippase ATP8A1-CDC50 but also provide insights into the intricate details of lipid transport.
Collapse
Affiliation(s)
- Song Yang
- Peking-Tsinghua Center for Life Sciences, Academy for Advanced Interdisciplinary Studies, Peking University, Beijing 100871, China
- Center for Quantitative Biology, Academy for Advanced Interdisciplinary Studies, Peking University, Beijing 100871, China
| | - Chen Song
- Peking-Tsinghua Center for Life Sciences, Academy for Advanced Interdisciplinary Studies, Peking University, Beijing 100871, China
- Center for Quantitative Biology, Academy for Advanced Interdisciplinary Studies, Peking University, Beijing 100871, China
| |
Collapse
|
31
|
Wu X, Lin H, Bai R, Duan H. Deep learning for advancing peptide drug development: Tools and methods in structure prediction and design. Eur J Med Chem 2024; 268:116262. [PMID: 38387334 DOI: 10.1016/j.ejmech.2024.116262] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/04/2024] [Revised: 02/06/2024] [Accepted: 02/17/2024] [Indexed: 02/24/2024]
Abstract
Peptides can bind challenging disease targets with high affinity and specificity, offering enormous opportunities for addressing unmet medical needs. However, peptides' unique features, including smaller size, increased structural flexibility, and limited data availability, pose additional challenges to the design process compared to proteins. This review explores the dynamic field of peptide therapeutics, leveraging deep learning to enhance structure prediction and design. Our exploration encompasses various facets of peptide research, ranging from dataset curation handling to model development. As deep learning technologies become more refined, we channel our efforts into peptide structure prediction and design, aligning with the fundamental principles of structure-activity relationships in drug development. To guide researchers in harnessing the potential of deep learning to advance peptide drug development, our insights comprehensively explore current challenges and future directions of peptide therapeutics.
Collapse
Affiliation(s)
- Xinyi Wu
- College of Pharmaceutical Sciences, Zhejiang University of Technology, Hangzhou, 310014, PR China
| | - Huitian Lin
- College of Pharmaceutical Sciences, Zhejiang University of Technology, Hangzhou, 310014, PR China
| | - Renren Bai
- School of Pharmacy, Hangzhou Normal University, Hangzhou, 311121, PR China.
| | - Hongliang Duan
- Faculty of Applied Sciences, Macao Polytechnic University, Macao, 999078, PR China.
| |
Collapse
|
32
|
Tu G, Fu T, Zheng G, Xu B, Gou R, Luo D, Wang P, Xue W. Computational Chemistry in Structure-Based Solute Carrier Transporter Drug Design: Recent Advances and Future Perspectives. J Chem Inf Model 2024; 64:1433-1455. [PMID: 38294194 DOI: 10.1021/acs.jcim.3c01736] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 02/01/2024]
Abstract
Solute carrier transporters (SLCs) are a class of important transmembrane proteins that are involved in the transportation of diverse solute ions and small molecules into cells. There are approximately 450 SLCs within the human body, and more than a quarter of them are emerging as attractive therapeutic targets for multiple complex diseases, e.g., depression, cancer, and diabetes. However, only 44 unique transporters (∼9.8% of the SLC superfamily) with 3D structures and specific binding sites have been reported. To design innovative and effective drugs targeting diverse SLCs, there are a number of obstacles that need to be overcome. However, computational chemistry, including physics-based molecular modeling and machine learning- and deep learning-based artificial intelligence (AI), provides an alternative and complementary way to the classical drug discovery approach. Here, we present a comprehensive overview on recent advances and existing challenges of the computational techniques in structure-based drug design of SLCs from three main aspects: (i) characterizing multiple conformations of the proteins during the functional process of transportation, (ii) identifying druggability sites especially the cryptic allosteric ones on the transporters for substrates and drugs binding, and (iii) discovering diverse small molecules or synthetic protein binders targeting the binding sites. This work is expected to provide guidelines for a deep understanding of the structure and function of the SLC superfamily to facilitate rational design of novel modulators of the transporters with the aid of state-of-the-art computational chemistry technologies including artificial intelligence.
Collapse
Affiliation(s)
- Gao Tu
- Chongqing Key Laboratory of Natural Product Synthesis and Drug Research, School of Pharmaceutical Sciences, Chongqing University, Chongqing 401331, China
| | - Tingting Fu
- College of Pharmaceutical Sciences, Zhejiang University, Hangzhou 310058, China
| | | | - Binbin Xu
- Chengdu Sintanovo Biotechnology Co., Ltd., Chengdu 610200, China
| | - Rongpei Gou
- Chongqing Key Laboratory of Natural Product Synthesis and Drug Research, School of Pharmaceutical Sciences, Chongqing University, Chongqing 401331, China
| | - Ding Luo
- Chongqing Key Laboratory of Natural Product Synthesis and Drug Research, School of Pharmaceutical Sciences, Chongqing University, Chongqing 401331, China
| | - Panpan Wang
- College of Chemistry and Pharmaceutical Engineering, Huanghuai University, Zhumadian 463000, China
| | - Weiwei Xue
- Chongqing Key Laboratory of Natural Product Synthesis and Drug Research, School of Pharmaceutical Sciences, Chongqing University, Chongqing 401331, China
| |
Collapse
|
33
|
Mróz J, Pelc M, Mitusińska K, Chorostowska-Wynimko J, Jezela-Stanek A. Computational Tools to Assist in Analyzing Effects of the SERPINA1 Gene Variation on Alpha-1 Antitrypsin (AAT). Genes (Basel) 2024; 15:340. [PMID: 38540399 PMCID: PMC10970068 DOI: 10.3390/genes15030340] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/13/2024] [Revised: 02/28/2024] [Accepted: 03/04/2024] [Indexed: 06/14/2024] Open
Abstract
In the rapidly advancing field of bioinformatics, the development and application of computational tools to predict the effects of single nucleotide variants (SNVs) are shedding light on the molecular mechanisms underlying disorders. Also, they hold promise for guiding therapeutic interventions and personalized medicine strategies in the future. A comprehensive understanding of the impact of SNVs in the SERPINA1 gene on alpha-1 antitrypsin (AAT) protein structure and function requires integrating bioinformatic approaches. Here, we provide a guide for clinicians to navigate through the field of computational analyses which can be applied to describe a novel genetic variant. Predicting the clinical significance of SERPINA1 variation allows clinicians to tailor treatment options for individuals with alpha-1 antitrypsin deficiency (AATD) and related conditions, ultimately improving the patient's outcome and quality of life. This paper explores the various bioinformatic methodologies and cutting-edge approaches dedicated to the assessment of molecular variants of genes and their product proteins using SERPINA1 and AAT as an example.
Collapse
Affiliation(s)
- Jakub Mróz
- Tunneling Group, Biotechnology Center, Silesian University of Technology, Krzywoustego St. 8, 44-100 Gliwice, Poland;
| | - Magdalena Pelc
- Department of Genetics and Clinical Immunology, National Institute of Tuberculosis and Lung Diseases, 26 Plocka St., 01-138 Warsaw, Poland; (M.P.); (J.C.-W.)
| | - Karolina Mitusińska
- Tunneling Group, Biotechnology Center, Silesian University of Technology, Krzywoustego St. 8, 44-100 Gliwice, Poland;
| | - Joanna Chorostowska-Wynimko
- Department of Genetics and Clinical Immunology, National Institute of Tuberculosis and Lung Diseases, 26 Plocka St., 01-138 Warsaw, Poland; (M.P.); (J.C.-W.)
| | - Aleksandra Jezela-Stanek
- Department of Genetics and Clinical Immunology, National Institute of Tuberculosis and Lung Diseases, 26 Plocka St., 01-138 Warsaw, Poland; (M.P.); (J.C.-W.)
| |
Collapse
|
34
|
Jänes J, Beltrao P. Deep learning for protein structure prediction and design-progress and applications. Mol Syst Biol 2024; 20:162-169. [PMID: 38291232 PMCID: PMC10912668 DOI: 10.1038/s44320-024-00016-x] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 06/26/2023] [Revised: 12/21/2023] [Accepted: 01/11/2024] [Indexed: 02/01/2024] Open
Abstract
Proteins are the key molecular machines that orchestrate all biological processes of the cell. Most proteins fold into three-dimensional shapes that are critical for their function. Studying the 3D shape of proteins can inform us of the mechanisms that underlie biological processes in living cells and can have practical applications in the study of disease mutations or the discovery of novel drug treatments. Here, we review the progress made in sequence-based prediction of protein structures with a focus on applications that go beyond the prediction of single monomer structures. This includes the application of deep learning methods for the prediction of structures of protein complexes, different conformations, the evolution of protein structures and the application of these methods to protein design. These developments create new opportunities for research that will have impact across many areas of biomedical research.
Collapse
Affiliation(s)
- Jürgen Jänes
- Institute of Molecular Systems Biology, ETH Zürich, 8093, Zürich, Switzerland
- Swiss Institute of Bioinformatics, Lausanne, Switzerland
| | - Pedro Beltrao
- Institute of Molecular Systems Biology, ETH Zürich, 8093, Zürich, Switzerland.
- Swiss Institute of Bioinformatics, Lausanne, Switzerland.
| |
Collapse
|
35
|
Nam K, Shao Y, Major DT, Wolf-Watz M. Perspectives on Computational Enzyme Modeling: From Mechanisms to Design and Drug Development. ACS OMEGA 2024; 9:7393-7412. [PMID: 38405524 PMCID: PMC10883025 DOI: 10.1021/acsomega.3c09084] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Grants] [Track Full Text] [Download PDF] [Figures] [Subscribe] [Scholar Register] [Received: 11/14/2023] [Revised: 01/15/2024] [Accepted: 01/19/2024] [Indexed: 02/27/2024]
Abstract
Understanding enzyme mechanisms is essential for unraveling the complex molecular machinery of life. In this review, we survey the field of computational enzymology, highlighting key principles governing enzyme mechanisms and discussing ongoing challenges and promising advances. Over the years, computer simulations have become indispensable in the study of enzyme mechanisms, with the integration of experimental and computational exploration now established as a holistic approach to gain deep insights into enzymatic catalysis. Numerous studies have demonstrated the power of computer simulations in characterizing reaction pathways, transition states, substrate selectivity, product distribution, and dynamic conformational changes for various enzymes. Nevertheless, significant challenges remain in investigating the mechanisms of complex multistep reactions, large-scale conformational changes, and allosteric regulation. Beyond mechanistic studies, computational enzyme modeling has emerged as an essential tool for computer-aided enzyme design and the rational discovery of covalent drugs for targeted therapies. Overall, enzyme design/engineering and covalent drug development can greatly benefit from our understanding of the detailed mechanisms of enzymes, such as protein dynamics, entropy contributions, and allostery, as revealed by computational studies. Such a convergence of different research approaches is expected to continue, creating synergies in enzyme research. This review, by outlining the ever-expanding field of enzyme research, aims to provide guidance for future research directions and facilitate new developments in this important and evolving field.
Collapse
Affiliation(s)
- Kwangho Nam
- Department
of Chemistry and Biochemistry, University
of Texas at Arlington, Arlington, Texas 76019, United States
| | - Yihan Shao
- Department
of Chemistry and Biochemistry, University
of Oklahoma, Norman, Oklahoma 73019-5251, United States
| | - Dan T. Major
- Department
of Chemistry and Institute for Nanotechnology & Advanced Materials, Bar-Ilan University, Ramat-Gan 52900, Israel
| | | |
Collapse
|
36
|
Raisinghani N, Alshahrani M, Gupta G, Tian H, Xiao S, Tao P, Verkhivker G. Interpretable Atomistic Prediction and Functional Analysis of Conformational Ensembles and Allosteric States in Protein Kinases Using AlphaFold2 Adaptation with Randomized Sequence Scanning and Local Frustration Profiling. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.02.15.580591. [PMID: 38496487 PMCID: PMC10942451 DOI: 10.1101/2024.02.15.580591] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 03/19/2024]
Abstract
The groundbreaking achievements of AlphaFold2 (AF2) approaches in protein structure modeling marked a transformative era in structural biology. Despite the success of AF2 tools in predicting single protein structures, these methods showed intrinsic limitations in predicting multiple functional conformations of allosteric proteins and fold-switching systems. The recent NMR-based structural determination of the unbound ABL kinase in the active state and two inactive low-populated functional conformations that are unique for ABL kinase presents an ideal challenge for AF2 approaches. In the current study we employ several implementations of AF2 methods to predict protein conformational ensembles and allosteric states of the ABL kinase including (a) multiple sequence alignments (MSA) subsampling approach; (b) SPEACH_AF approach in which alanine scanning is performed on generated MSAs; and (c) introduced in this study randomized full sequence mutational scanning for manipulation of sequence variations combined with the MSA subsampling. We show that the proposed AF2 adaptation combined with local frustration mapping of conformational states enable accurate prediction of the ABL active and intermediate structures and conformational ensembles, also offering a robust approach for interpretable characterization of the AF2 predictions and limitations in detecting hidden allosteric states. We found that the large high frustration residue clusters are uniquely characteristic of the low-populated, fully inactive ABL form and can define energetically frustrated cracking sites of conformational transitions, presenting difficult targets for AF2 methods. This study uncovered previously unappreciated, fundamental connections between distinct patterns of local frustration in functional kinase states and AF2 successes/limitations in detecting low-populated frustrated conformations, providing a better understanding of benefits and limitations of current AF2-based adaptations in modeling of conformational ensembles.
Collapse
|
37
|
Chu AE, Lu T, Huang PS. Sparks of function by de novo protein design. Nat Biotechnol 2024; 42:203-215. [PMID: 38361073 DOI: 10.1038/s41587-024-02133-2] [Citation(s) in RCA: 1] [Impact Index Per Article: 1.0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 11/05/2023] [Accepted: 01/09/2024] [Indexed: 02/17/2024]
Abstract
Information in proteins flows from sequence to structure to function, with each step causally driven by the preceding one. Protein design is founded on inverting this process: specify a desired function, design a structure executing this function, and find a sequence that folds into this structure. This 'central dogma' underlies nearly all de novo protein-design efforts. Our ability to accomplish these tasks depends on our understanding of protein folding and function and our ability to capture this understanding in computational methods. In recent years, deep learning-derived approaches for efficient and accurate structure modeling and enrichment of successful designs have enabled progression beyond the design of protein structures and towards the design of functional proteins. We examine these advances in the broader context of classical de novo protein design and consider implications for future challenges to come, including fundamental capabilities such as sequence and structure co-design and conformational control considering flexibility, and functional objectives such as antibody and enzyme design.
Collapse
Affiliation(s)
- Alexander E Chu
- Biophysics Program, Stanford University, Palo Alto, CA, USA
- Department of Bioengineering, Stanford University, Palo Alto, CA, USA
- Google DeepMind, London, UK
| | - Tianyu Lu
- Department of Bioengineering, Stanford University, Palo Alto, CA, USA
| | - Po-Ssu Huang
- Biophysics Program, Stanford University, Palo Alto, CA, USA.
- Department of Bioengineering, Stanford University, Palo Alto, CA, USA.
| |
Collapse
|
38
|
Licht JA, Berry SP, Gutierrez MA, Gaudet R. They all rock: A systematic comparison of conformational movements in LeuT-fold transporters. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.01.24.577062. [PMID: 38352416 PMCID: PMC10862720 DOI: 10.1101/2024.01.24.577062] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Figures] [Subscribe] [Scholar Register] [Indexed: 02/19/2024]
Abstract
Many membrane transporters share the LeuT fold-two five-helix repeats inverted across the membrane plane. Despite hundreds of structures, whether distinct conformational mechanisms are supported by the LeuT fold has not been systematically determined. After annotating published LeuT-fold structures, we analyzed distance difference matrices (DDMs) for nine proteins with multiple available conformations. We identified rigid bodies and relative movements of transmembrane helices (TMs) during distinct steps of the transport cycle. In all transporters the bundle (first two TMs of each repeat) rotates relative to the hash (third and fourth TMs). Motions of the arms (fifth TM) to close or open the intracellular and outer vestibules are common, as is a TM1a swing, with notable variations in the opening-closing motions of the outer vestibule. Our analyses suggest that LeuT-fold transporters layer distinct motions on a common bundle-hash rock and demonstrate that systematic analyses can provide new insights into large structural datasets.
Collapse
Affiliation(s)
- Jacob A. Licht
- Department of Molecular and Cellular Biology, Harvard University, Cambridge, MA, USA
| | - Samuel P. Berry
- Department of Molecular and Cellular Biology, Harvard University, Cambridge, MA, USA
| | - Michael A. Gutierrez
- Department of Molecular and Cellular Biology, Harvard University, Cambridge, MA, USA
- Present address: Novartis Biomedical Research, Cambridge, MA, USA
| | - Rachelle Gaudet
- Department of Molecular and Cellular Biology, Harvard University, Cambridge, MA, USA
- Lead contact
| |
Collapse
|
39
|
Ohnuki J, Jaunet-Lahary T, Yamashita A, Okazaki KI. Accelerated Molecular Dynamics and AlphaFold Uncover a Missing Conformational State of Transporter Protein OxlT. J Phys Chem Lett 2024; 15:725-732. [PMID: 38215403 DOI: 10.1021/acs.jpclett.3c03052] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 01/14/2024]
Abstract
Transporter proteins change their conformations to carry their substrate across the cell membrane. The conformational dynamics is vital to understanding the transport function. We have studied the oxalate transporter (OxlT), an oxalate:formate antiporter from Oxalobacter formigenes, significant in avoiding kidney stone formation. The atomic structure of OxlT has been recently solved in the outward-open and occluded states. However, the inward-open conformation is still missing, hindering a complete understanding of the transporter. Here, we performed a Gaussian accelerated molecular dynamics simulation to sample the extensive conformational space of OxlT and successfully predicted the inward-open conformation where cytoplasmic substrate formate binding was preferred over oxalate binding. We also identified critical interactions for the inward-open conformation. The results were complemented by an AlphaFold2 structure prediction. Although AlphaFold2 solely predicted OxlT in the outward-open conformation, mutation of the identified critical residues made it partly predict the inward-open conformation, identifying possible state-shifting mutations.
Collapse
Affiliation(s)
- Jun Ohnuki
- Research Center for Computational Science, Institute for Molecular Science, National Institutes of Natural Sciences, Okazaki 444-8585, Japan
- Graduate Institute for Advanced Studies, SOKENDAI, Okazaki, Aichi 444-8585, Japan
| | - Titouan Jaunet-Lahary
- Research Center for Computational Science, Institute for Molecular Science, National Institutes of Natural Sciences, Okazaki 444-8585, Japan
| | - Atsuko Yamashita
- Graduate School of Medicine, Dentistry and Pharmaceutical Sciences, Okayama University, Okayama 700-8530, Japan
| | - Kei-Ichi Okazaki
- Research Center for Computational Science, Institute for Molecular Science, National Institutes of Natural Sciences, Okazaki 444-8585, Japan
- Graduate Institute for Advanced Studies, SOKENDAI, Okazaki, Aichi 444-8585, Japan
| |
Collapse
|
40
|
Yan Z, Wang J. Evolution shapes interaction patterns for epistasis and specific protein binding in a two-component signaling system. Commun Chem 2024; 7:13. [PMID: 38233668 PMCID: PMC10794238 DOI: 10.1038/s42004-024-01098-2] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 06/27/2023] [Accepted: 01/05/2024] [Indexed: 01/19/2024] Open
Abstract
The elegant design of protein sequence/structure/function relationships arises from the interaction patterns between amino acid positions. A central question is how evolutionary forces shape the interaction patterns that encode long-range epistasis and binding specificity. Here, we combined family-wide evolutionary analysis of natural homologous sequences and structure-oriented evolution simulation for two-component signaling (TCS) system. The magnitude-frequency relationship of coupling conservation between positions manifests a power-law-like distribution and the positions with highly coupling conservation are sparse but distributed intensely on the binding surfaces and hydrophobic core. The structure-specific interaction pattern involves further optimization of local frustrations at or near the binding surface to adapt the binding partner. The construction of family-wide conserved interaction patterns and structure-specific ones demonstrates that binding specificity is modulated by both direct intermolecular interactions and long-range epistasis across the binding complex. Evolution sculpts the interaction patterns via sequence variations at both family-wide and structure-specific levels for TCS system.
Collapse
Affiliation(s)
- Zhiqiang Yan
- Center for Theoretical Interdisciplinary Sciences, Wenzhou Institute, University of Chinese Academy of Sciences, Wenzhou, Zhejiang, 325001, PR China
| | - Jin Wang
- Department of Chemistry and Physics, State University of New York at Stony Brook, Stony Brook, NY, 11790, USA.
| |
Collapse
|
41
|
Schafer JW, Chakravarty D, Chen EA, Porter LL. Sequence clustering confounds AlphaFold2. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.01.05.574434. [PMID: 38313252 PMCID: PMC10836070 DOI: 10.1101/2024.01.05.574434] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 02/06/2024]
Abstract
Though typically associated with a single folded state, some globular proteins remodel their secondary and/or tertiary structures in response to cellular stimuli. AlphaFold21 (AF2) readily generates one dominant protein structure for these fold-switching (a.k.a. metamorphic) proteins2, but it often fails to predict their alternative experimentally observed structures3,4. Wayment-Steele, et al. steered AF2 to predict alternative structures of a few metamorphic proteins using a method they call AF-cluster5. However, their Paper lacks some essential controls needed to assess AF-cluster's reliability. We find that these controls show AF-cluster to be a poor predictor of metamorphic proteins. First, closer examination of the Paper's results reveals that random sequence sampling outperforms sequence clustering, challenging the claim that AF-cluster works by "deconvolving conflicting sets of couplings." Further, we observe that AF-cluster mistakes some single-folding KaiB homologs for fold switchers, a critical flaw bound to mislead users. Finally, proper error analysis reveals that AF-cluster predicts many correct structures with low confidence and some experimentally unobserved conformations with confidences similar to experimentally observed ones. For these reasons, we suggest using ColabFold6-based random sequence sampling7-augmented by other predictive approaches-as a more accurate and less computationally intense alternative to AF-cluster.
Collapse
Affiliation(s)
- Joseph W. Schafer
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Devlina Chakravarty
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Ethan A. Chen
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Lauren L. Porter
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
- Biochemistry and Biophysics Center, National Heart, Lung, and Blood Institute, National Institutes of Health, Bethesda, MD, 20892
| |
Collapse
|
42
|
Chakravarty D, Schafer JW, Chen EA, Thole JR, Porter LL. AlphaFold2 has more to learn about protein energy landscapes. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2023:2023.12.12.571380. [PMID: 38168383 PMCID: PMC10760193 DOI: 10.1101/2023.12.12.571380] [Citation(s) in RCA: 8] [Impact Index Per Article: 8.0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 01/05/2024]
Abstract
Recent work suggests that AlphaFold2 (AF2)-a deep learning-based model that can accurately infer protein structure from sequence-may discern important features of folded protein energy landscapes, defined by the diversity and frequency of different conformations in the folded state. Here, we test the limits of its predictive power on fold-switching proteins, which assume two structures with regions of distinct secondary and/or tertiary structure. Using several implementations of AF2, including two published enhanced sampling approaches, we generated >280,000 models of 93 fold-switching proteins whose experimentally determined conformations were likely in AF2's training set. Combining all models, AF2 predicted fold switching with a modest success rate of ~25%, indicating that it does not readily sample both experimentally characterized conformations of most fold switchers. Further, AF2's confidence metrics selected against models consistent with experimentally determined fold-switching conformations in favor of inconsistent models. Accordingly, these confidence metrics-though suggested to evaluate protein energetics reliably-did not discriminate between low and high energy states of fold-switching proteins. We then evaluated AF2's performance on seven fold-switching proteins outside of its training set, generating >159,000 models in total. Fold switching was accurately predicted in one of seven targets with moderate confidence. Further, AF2 demonstrated no ability to predict alternative conformations of two newly discovered targets without homologs in the set of 93 fold switchers. These results indicate that AF2 has more to learn about the underlying energetics of protein ensembles and highlight the need for further developments of methods that readily predict multiple protein conformations.
Collapse
Affiliation(s)
- Devlina Chakravarty
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Joseph W. Schafer
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Ethan A. Chen
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Joseph R. Thole
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
- Biochemistry and Biophysics Center, National Heart, Lung, and Blood Institute, National Institutes of Health, Bethesda, MD, 20892
| | - Lauren L. Porter
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
- Biochemistry and Biophysics Center, National Heart, Lung, and Blood Institute, National Institutes of Health, Bethesda, MD, 20892
| |
Collapse
|
43
|
Raisinghani N, Alshahrani M, Gupta G, Xiao S, Tao P, Verkhivker G. AlphaFold2-Enabled Atomistic Modeling of Epistatic Binding Mechanisms for the SARS-CoV-2 Spike Omicron XBB.1.5, EG.5 and FLip Variants: Convergent Evolution Hotspots Cooperate to Control Stability and Conformational Adaptability in Balancing ACE2 Binding and Antibody Resistance. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2023:2023.12.11.571185. [PMID: 38168257 PMCID: PMC10760024 DOI: 10.1101/2023.12.11.571185] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 01/05/2024]
Abstract
In this study, we combined AI-based atomistic structural modeling and microsecond molecular simulations of the SARS-CoV-2 Spike complexes with the host receptor ACE2 for XBB.1.5+L455F, XBB.1.5+F456L(EG.5) and XBB.1.5+L455F/F456L (FLip) lineages to examine the mechanisms underlying the role of convergent evolution hotspots in balancing ACE2 binding and antibody evasion. Using the ensemble-based mutational scanning of the spike protein residues and physics-based rigorous computations of binding affinities, we identified binding energy hotspots and characterized molecular basis underlying epistatic couplings between convergent mutational hotspots. Consistent with the experiments, the results revealed the mediating role of Q493 hotspot in synchronization of epistatic couplings between L455F and F456L mutations providing a quantitative insight into the mechanism underlying differences between XBB lineages. Mutational profiling is combined with network-based model of epistatic couplings showing that the Q493, L455 and F456 sites mediate stable communities at the binding interface with ACE2 and can serve as stable mediators of non-additive couplings. Structure-based mutational analysis of Spike protein binding with the class 1 antibodies quantified the critical role of F456L and F486P mutations in eliciting strong immune evasion response. The results of this analysis support a mechanism in which the emergence of EG.5 and FLip variants may have been dictated by leveraging strong epistatic effects between several convergent revolutionary hotspots that provide synergy between the improved ACE2 binding and broad neutralization resistance. This interpretation is consistent with the notion that functionally balanced substitutions which simultaneously optimize immune evasion and high ACE2 affinity may continue to emerge through lineages with beneficial pair or triplet combinations of RBD mutations involving mediators of epistatic couplings and sites in highly adaptable RBD regions.
Collapse
|
44
|
Osman S. Space exploration: finding new protein conformations using AlphaFold2. Nat Struct Mol Biol 2023; 30:1835. [PMID: 38087089 DOI: 10.1038/s41594-023-01186-2] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 12/18/2023]
Affiliation(s)
- Sara Osman
- Nature Structural & Molecular Biology, .
| |
Collapse
|
45
|
Porter LL, Chakravarty D, Schafer JW, Chen EA. ColabFold predicts alternative protein structures from single sequences, coevolution unnecessary for AF-cluster. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2023:2023.11.21.567977. [PMID: 38076792 PMCID: PMC10705582 DOI: 10.1101/2023.11.21.567977] [Citation(s) in RCA: 6] [Impact Index Per Article: 6.0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Figures] [Subscribe] [Scholar Register] [Indexed: 12/22/2023]
Abstract
Though typically associated with a single folded state, globular proteins are dynamic and often assume alternative or transient structures important for their functions1,2. Wayment-Steele, et al. steered ColabFold3 to predict alternative structures of several proteins using a method they call AF-cluster4. They propose that AF-cluster "enables ColabFold to sample alternate states of known metamorphic proteins with high confidence" by first clustering multiple sequence alignments (MSAs) in a way that "deconvolves" coevolutionary information specific to different conformations and then using these clusters as input for ColabFold. Contrary to this Coevolution Assumption, clustered MSAs are not needed to make these predictions. Rather, these alternative structures can be predicted from single sequences and/or sequence similarity, indicating that coevolutionary information is unnecessary for predictive success and may not be used at all. These results suggest that AF-cluster's predictive scope is likely limited to sequences with distinct-yet-homologous structures within ColabFold's training set.
Collapse
Affiliation(s)
- Lauren L. Porter
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
- Biochemistry and Biophysics Center, National Heart, Lung, and Blood Institute, National Institutes of Health, Bethesda, MD, 20892
| | - Devlina Chakravarty
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Joseph W. Schafer
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| | - Ethan A. Chen
- National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894
| |
Collapse
|