Document text
Principal Investigator: Vipul Periwal
Organization: NATIONAL INSTITUTE OF DIABETES AND DIGESTIVE AND KIDNEY DISEASES
Fiscal Year: 2024
Award: $189,847
Funding agency: National Institute of Diabetes and Digestive and Kidney Diseases
Current models for inferring co-evolutionary contacts from aligned protein sequence data use the principle of maximum-entropy (ME) which assumes that the sequences are in a state of equilibrium with respect to the evolutionary forces acting on them. This has two obvious issues. Firstly, HIV protein sequences cannot accurately be defined as being in a state of equilibrium, especially when our aim is to discover possible dynamic mutation-driven escape pathways. Secondly, epistatic interactions are not necessarily symmetric because proteins do not exist in isolation but must interact with other biological macromolecules. Thus, interacting residues are not symmetrically equal because, for instance, one might be structural and the other functional in their respective contributions to the proteins function.
To overcome these limitations, we developed a methodology which seeks to describe the same physical system, but in a non-equilibrium state. This method, termed Expectation Reflection (ER), has been shown to improve upon the maximum-entropy approach and successfully characterize epistatic interactions within the SARS-CoV-2 genome. The ER approach directly computes the conditional probability of observing a mutated amino acid at a position, given a sequence in the context of a population of sequences. With this conditional probability we can directly calculate the likelihood of a specific residue, such as a DRM, in the context of a given population. The conditional probability also means that our inferred pair-wise interactions are not necessarily symmetric. Ultimately this methodology gives a more tractable, and biologically appropriate, theoretical description of the transition energy of a given mutation which is the basis of several analyses of DRM in previous work.
Dr. Kearney has indicated to us that at this point the most valuable part of these analyses would be to help in characterizing broadly neutralizing antibodies because the success of the anti-retroviral treatment regimen in people who comply with the drug regimen implies that in these individuals the mutation rate of the virus does not overwhelm the immune system. In other words, the therapeutic focus needs to be shifted to the development and characterization of broadly neutralizing antibodies in the population that is not adequately protected by the drug regimen. Thus we are focusing our work on the viral envelope glycoprotein Env.
Sequence covariation in multiple sequence alignments of homologous proteins has been used extensively to obtain insights into protein structure and function. We present a novel method for finding coevolutionary relationships within protein sequences using an iterative influence estimator that infers the influence of the rest of the protein sequence on the presence or absence of a specific residue at a particular position in the amino-acid sequence. This mechanistic approach yields a directed coupling matrix that is asymmetric in contrast to previous energy-based approaches and allows us to expand beyond structure prediction of proteins to characterize the likelihood of specific future mutations. We applied this technique as an extension of a similar analysis on sequence evolution of the active lineage of LINE-1 retrotransposons. We find that some mutation trajectories show large deviations and escape to chronologically adjacent LINE-1 families, predicting evolutionary pathways. Such pathways require mutations that are probabilistically rare. We used our model to compute the probabilities of these trajectories and determine the ones that are least unlikely, which we define as large deviation evolutionary pathways. As another important application of our method, we analyzed entrenchment of drug resistance in HIV-1 subtype B. Primary mutations in HIV that cause drug resistance might be unfavorable in the wild-type background but can become favorable (or entrenched) when accompanied with specific secondary mutations. Our method revealed complex patterns of secondary mutations that can make drug resistant mutations highly favorable. Lastly, we applied our method to millions of sequences from the Phylogenetic Assignment of Named Global Outbreak (PANGO) lineages of different SARS-COV2 proteins to determine large deviation evolutionary pathways between different strains of the virus.
Terms: <2019 novel corona virus><2019 novel coronavirus><2019-nCoV><AIDS Virus><Accounting><Acquired Immune Deficiency Syndrome Virus><Acquired Immunodeficiency Syndrome Virus><Amino Acid Sequence><Amino Acids><Biological><COVID-19 genome><COVID-19 virus><COVID-19 virus genome><COVID19 genome><COVID19 virus><COVID19 virus genome><Capsid><Chronology><CoV-2><CoV2><Complex><Coupling><Data><Data Collection><Data Set><Data Storage and Retrieval><Development><Disease Outbreaks><Drug Therapy><Drug resistance><Drugs><EC 2.7.7.49><Entropy><Epistasis><Epistatic Deviation><Equilibrium><Esteroproteases><Evolution><Face><Family><Future><Generalized Growth><Genetic Alteration><Genetic Anticipation><Genetic Change><Genetic Epistasis><Genetic defect><Growth><HIV><HIV 1 drug resistance><HIV-1 drug resistance><HIV-1 drug resistant><HIV1 drug resistance><HIV1 drug resistant><Homologous Protein><Human Immunodeficiency Viruses><Immune system><Individual><Infection><Information Sciences><Integrase><Interaction Deviation><LAV-HTLV-III><Lead><Lymphadenopathy-Associated Virus><Medication><Methodology><Methods><Modeling><Mutate><Mutation><Names><Outbreaks><Pathway interactions><Pattern><Pb element><Peptidases><Peptide Hydrolases><Persons><Pharmaceutical Preparations><Pharmacotherapy><Phylogenetic Analysis><Phylogenetics><Population><Position><Positioning Attribute><Primary Protein Structure><Probability><Protease Gene><Proteases><Protein Homolog><ProteinHomolog><Proteinases><Proteins><Proteolytic Enzymes><Publications><RNA Transcriptase><RNA-Dependent DNA Polymerase><RNA-Directed DNA Polymerase><Regimen><Rest><Retrotransposon><Reverse Transcriptase><Revertase><Role><Route><SARS corona virus 2><SARS-CO-V2><SARS-COVID-2><SARS-CoV-2><SARS-CoV-2 genome><SARS-CoV2><SARS-CoV2 genome><SARS-associated corona virus 2><SARS-associated coronavirus 2><SARS-coronavirus-2><SARS-related corona virus 2><SARS-related coronavirus 2><SARSCoV2><SEQ-AN><Science><Scientific Publication><Sequence Alignment><Sequence Analyses><Sequence Analysis><Severe Acute Respiratory Coronavirus 2><Severe Acute Respiratory Distress Syndrome CoV 2><Severe Acute Respiratory Distress Syndrome Corona Virus 2><Severe Acute Respiratory Distress Syndrome Coronavirus 2><Severe Acute Respiratory Syndrome CoV 2><Severe Acute Respiratory Syndrome-associated coronavirus 2><Severe Acute Respiratory Syndrome-related coronavirus 2><Severe acute respiratory syndrome associated corona virus 2><Severe acute respiratory syndrome coronavirus 2><Severe acute respiratory syndrome related corona virus 2><System><Techniques><Technology><Therapeutic><Tissue Growth><Treatment Protocols><Treatment Regimen><Treatment Schedule><Viral><Viral Gene Products><Viral Gene Proteins><Viral Genome><Viral Proteins><Virus><Virus-HIV><Work><Wuhan coronavirus><aminoacid><antiretroviral therapy><antiretroviral treatment><balance><balance function><biologic><combinatorial><computational resources><computing resources><coronavirus disease 2019 genome><coronavirus disease 2019 virus><coronavirus disease 2019 virus genome><coronavirus disease-19 virus><data retrieval><data storage><developmental><drug resistance in HIV 1><drug resistance in HIV-1><drug resistant><drug resistant HIV 1><drug resistant HIV-1><drug treatment><drug/agent><dynamic mutation><env Glycoproteins><epistatic relationship><expectation><faces><facial><fitness><gene x gene interaction><genetic epistases><genome mutation><genome wide analysis><genome wide studies><genome-wide analysis><genome-wide identification><hCoV19><heavy metal Pb><heavy metal lead><heuristics><improved><insight><large data sets><large datasets><large scale data><large scale data sets><large scale datasets><lens><lenses><macromolecule><nCoV2><name><named><naming><neutralizing antibody><novel><ontogeny><pathogen><pathway><pressure><protein function><protein sequence><protein structure><protein structure prediction><protein structures><proteins structure><resistance mutation><resistance to Drug><resistance to HIV-1 drug><resistance to HIV1 drug><resistant mutation><resistant to Drug><resistant to HIV-1 drug><resistant to HIV1 drug><sequencing alignment><severe acute respiratory syndrome coronavirus 2 genome><social role><success><therapeutic target><tool><viral fitness><virus genome><virus protein>