Pattern Identification in Sequence Activity Data

NIH Pandemic-Era Grants

Pandemic Era Grants

2023

Document text

Principal Investigator: Vipul  Periwal
Organization: NATIONAL INSTITUTE OF DIABETES AND DIGESTIVE AND KIDNEY DISEASES
Fiscal Year: 2023
Award: $339,647
Funding agency: National Institute of Diabetes and Digestive and Kidney Diseases

Current models for inferring co-evolutionary contacts from aligned protein sequence data use the principle of maximum-entropy (ME) which assumes that the sequences are in a state of equilibrium with respect to the evolutionary forces acting on them. This has two obvious issues. Firstly, HIV protein sequences cannot accurately be defined as being in a state of equilibrium, especially when our aim is to discover possible dynamic mutation-driven escape pathways. Secondly, epistatic interactions are not necessarily symmetric because proteins do not exist in isolation but must interact with other biological macromolecules. Thus, interacting residues are not symmetrically equal because, for instance, one might be structural and the other functional in their respective contributions to the proteins function. 
To overcome these limitations, we developed a methodology which seeks to describe the same physical system, but in a non-equilibrium state. This method, termed Expectation Reflection (ER), has been shown to improve upon the maximum-entropy approach and successfully characterize epistatic interactions within the SARS-CoV-2 genome. The ER approach directly computes the conditional probability of observing a mutated amino acid at a position, given a sequence in the context of a population of sequences. With this conditional probability we can directly calculate the likelihood of a specific residue, such as a DRM, in the context of a given population. The conditional probability also means that our inferred pair-wise interactions are not necessarily symmetric. Ultimately this methodology gives a more tractable, and biologically appropriate, theoretical description of the transition energy of a given mutation which is the basis of several analyses of DRM in previous work.
Dr. Kearney has indicated to us that at this point the most valuable part of these analyses would be to help in characterizing broadly neutralizing antibodies because the success of the anti-retroviral treatment regimen in people who comply with the drug regimen implies that in these individuals the mutation rate of the virus does not overwhelm the immune system. In other words, the therapeutic focus needs to be shifted to the development and characterization of broadly neutralizing antibodies in the population that is not adequately protected by the drug regimen. Thus we are focusing our work on the viral envelope glycoprotein Env.

Terms: <2019 novel corona virus><2019 novel coronavirus><2019-nCoV><2019-nCoV S protein><2019-nCoV spike glycoprotein><2019-nCoV spike protein><AIDS Virus><Accounting><Acquired Immune Deficiency Syndrome Virus><Acquired Immunodeficiency Syndrome Virus><Amino Acid Sequence><Amino Acids><B.1.1.529><B.1.617.2><Biological><COVID-19 S protein><COVID-19 genome><COVID-19 spike glycoprotein><COVID-19 spike protein><COVID-19 virus><COVID-19 virus genome><COVID19 S protein><COVID19 genome><COVID19 spike glycoprotein><COVID19 spike protein><COVID19 virus><COVID19 virus genome><Capsid><CoV-2><CoV2><Coupling><Data><Data Collection><Data Set><Data Storage and Retrieval><Delta variant><Development><Drug Therapy><Drug resistance><Drugs><EC 2.7.7.49><Entropy><Epistasis><Epistatic Deviation><Equilibrium><Esteroproteases><Evolution><Face><Generalized Growth><Genetic Alteration><Genetic Anticipation><Genetic Change><Genetic Epistasis><Genetic defect><Growth><HIV><Human Immunodeficiency Viruses><Immune system><Individual><Infection><Information Sciences><Integrase><Interaction Deviation><LAV-HTLV-III><Lead><Lymphadenopathy-Associated Virus><Medication><Methodology><Methods><Modeling><Mutate><Mutation><Omicron variant><Pathway interactions><Pattern><Pb element><Peptidases><Peptide Hydrolases><Persons><Pharmaceutic Preparations><Pharmaceutical Agent><Pharmaceutical Preparations><Pharmaceuticals><Pharmacologic Substance><Pharmacological Substance><Pharmacotherapy><Population><Position><Positioning Attribute><Primary Protein Structure><Probability><Protease Gene><Proteases><Proteinases><Proteins><Proteolytic Enzymes><Publications><RNA Transcriptase><RNA-Dependent DNA Polymerase><RNA-Directed DNA Polymerase><Regimen><Reverse Transcriptase><Revertase><Role><Route><SARS corona virus 2><SARS-CO-V2><SARS-COVID-2><SARS-CoV-2><SARS-CoV-2 B.1.1.529><SARS-CoV-2 B.1.617.2><SARS-CoV-2 S protein><SARS-CoV-2 delta><SARS-CoV-2 genome><SARS-CoV-2 omicron><SARS-CoV-2 omicron variant><SARS-CoV-2 spike glycoprotein><SARS-CoV-2 spike protein><SARS-CoV2><SARS-CoV2 S protein><SARS-CoV2 genome><SARS-CoV2 spike glycoprotein><SARS-CoV2 spike protein><SARS-associated corona virus 2><SARS-associated coronavirus 2><SARS-coronavirus-2><SARS-related corona virus 2><SARS-related coronavirus 2><SARSCoV2><SEQ-AN><Science><Scientific Publication><Sequence Analyses><Sequence Analysis><Severe Acute Respiratory Coronavirus 2><Severe Acute Respiratory Distress Syndrome CoV 2><Severe Acute Respiratory Distress Syndrome Corona Virus 2><Severe Acute Respiratory Distress Syndrome Coronavirus 2><Severe Acute Respiratory Syndrome CoV 2><Severe Acute Respiratory Syndrome-associated coronavirus 2><Severe Acute Respiratory Syndrome-related coronavirus 2><Severe acute respiratory syndrome associated corona virus 2><Severe acute respiratory syndrome coronavirus 2><Severe acute respiratory syndrome coronavirus 2 S protein><Severe acute respiratory syndrome coronavirus 2 spike glycoprotein><Severe acute respiratory syndrome coronavirus 2 spike protein><Severe acute respiratory syndrome related corona virus 2><Structure><System><Techniques><Technology><Therapeutic><Tissue Growth><Treatment Protocols><Treatment Regimen><Treatment Schedule><Viral><Viral Gene Products><Viral Gene Proteins><Viral Genome><Viral Proteins><Virus><Virus-HIV><Work><Wuhan coronavirus><aminoacid><anti-retroviral therapy><anti-retroviral treatment><antiretroviral therapy><antiretroviral treatment><balance><balance function><biologic><combinatorial><computational resources><computing resources><coronavirus disease 2019 S protein><coronavirus disease 2019 genome><coronavirus disease 2019 spike glycoprotein><coronavirus disease 2019 spike protein><coronavirus disease 2019 virus><coronavirus disease 2019 virus genome><coronavirus disease-19 virus><data retrieval><data storage><design><designing><developmental><drug resistant><drug treatment><drug/agent><dynamic mutation><env Glycoproteins><epistatic relationship><expectation><faces><facial><fitness><gene x gene interaction><genetic epistases><genome mutation><genome wide analysis><genome wide studies><genome-wide analysis><genome-wide identification><hCoV19><heavy metal Pb><heavy metal lead><heuristics><improved><insight><large data sets><large datasets><large scale data><large scale data sets><large scale datasets><lens><lenses><macromolecule><nCoV2><neutralizing antibody><novel><omicron variant of COVID-19><omicron variant of SARS-CoV-2><ontogeny><pathogen><pathway><pharmaceutical><pressure><protein function><protein sequence><resistance to Drug><resistant to Drug><severe acute respiratory syndrome coronavirus 2 B.1.1.529><severe acute respiratory syndrome coronavirus 2 B.1.617.2><severe acute respiratory syndrome coronavirus 2 genome><social role><success><therapeutic target><tool><viral fitness><virus genome><virus protein>