Document text
Principal Investigator: leping li
Organization: NATIONAL INSTITUTE OF ENVIRONMENTAL HEALTH SCIENCES
Fiscal Year: 2021
Award: $1,040,862
Funding agency: National Institute of Environmental Health Sciences
Project 1: Predicting tumor response to drugs based on gene-expression biomarkers of sensitivity learned from cancer cell lines
Studies that characterize human cancer cell lines and evaluate their sensitivity to drugs provide valuable information about the therapeutic potential and the possible mechanisms of action of those drugs. The Genomics of Drug Sensitivity in Cancer (GDSC) Project has assayed the sensitivity of 987 cancer cell lines to 320 compounds in their phase 1 (GDSC1) assay and of an additional 809 cancer cell lines to 175 compounds (some of which were included in the GDSC1 assay) in their phase 2 (GDSC2) assay. The sensitivity of each cancer cell line to the drugs was represented as an IC50 value (the concentration at which a cell line exhibited an absolute inhibition in growth of 50%; lower IC50 implies higher sensitivity). GDSC also quantified the basal level (without exposure to drug) gene expression of many of the cancer cell lines using microarray. Concomitantly, other consortia such as the CCLE (cancer cell line encyclopedia) also profiled genome-wide gene expression of many of the cancer cell lines using RNA-seq.
We use GA/KNN to build k-nearest neighbors predictive models for 453 drugs using data on gene expression and drug sensitivity (IC50) from cancer cell lines. We identified many known drug-gene interactions and uncovered several potentially novel drug-gene associations. Importantly, we further applied these predictive models to 17,000 bulk RNA-seq samples from TCGA and the GTEx database to predict drug sensitivity for both normal and tumor tissues. We created a web site for users to visualize and download our predicted data (https://manticore.niehs.nih.gov/cancerRxTissue). Using trametinib as an example, we showed that our approach can faithfully recapitulate the known tumor specificity of the drug. Our work, however, differs from the previous work in several ways: a) our analysis is more comprehensive by including the latest drug sensitivity data from GDSC2 for 453 drugs; b) our work emphasizes identification of putative biomarkers of sensitivity to drugs and potential therapeutic options for cancer subpopulations; and c) we also predict toxicity of drugs to normal tissues using transcriptomic data from normal human tissues available from both TCGA and GTEx project. If validated, our predictions could have clinical relevance for patients care. The manuscript is currently under reversion (BMC Genomics).
Project 2: Identifying expression biomarkers that are predictive of body mass index (BMI) or diabetic status
More than half of the US population is either overweight or obese, and rates are steadily rising, both in the US and globally. Given the increasing prevalence of obesity, it is important to understand how overweight individuals respond to chemical exposures, i.e., to see if their responses to exposures differ from those of individuals with normal weight. As a first step in addressing this question, Dr. Alison Harrill at DNTP, employed DO mice as a population model for human variability in metabolic diseases associated with consumption of a high-fat diet (HFD) without additional chemical exposures. In this study, DO mice consumed control diet (10% kcal from fat; N=75) or HFD (60% kcal from fat; N=75) for 13 weeks. As expected, at study end, animals fed an HFD on average gained more weight, had higher fasting blood glucose, and had impaired glucose tolerance compared to animals on the control diet. There was, however, a high degree of variability within dietary groups for many endpoints (e.g., fasting blood glucose, insulin, leptin) and a lack of correlation among them. These observations may indicate the presence of subgroups of mice with distinct metabolic profiles. To identify possible metabolic subtypes, Alisons group carried out genome-wide profiling of gene expression patterns of liver, fat and muscle of the two groups of DO mice.
To identify transcriptomic features that are associated with some of the key endpoints, we are in the process of applying a tree-based algorithm to each of the three tissue-specific RNA-seq datasets separately. Initially, we plan to focus on the liver dataset. We are particularly interested in the following four key endpoints assessed at the end of the study: (1) body weight gain; (2) fasting blood glucose; (3) cumulative serum glucose level measured as AUC/mg (area under the curve); (4) percentage change in leptin. We have completed the analyses and a manuscript for the work has been drafted.
Project 3: Mining electronic health care records for clinical features associated with the severity of COVID-19 infection.
SARS-CoV-2 (Covid-19) is a beta coronavirus that uses the angiotensin-converting enzyme 2 (ACE2) receptor to gain entry to host cells. Currently, no effective treatments for SARS-CoV-2 (COVID-19) are known, although several clinical trials are currently underway. The clinical manifestation for COVID-19 infection is highly variable, ranging from asymptomatic to fatal. The drivers of this marked variability remain largely unclear. Understanding the association between the clinical features and the severity of COVID-19 infection is critical for COVID-19 disease management and outcome improvement. Although several putative (bio)markers such as inflammatory cytokines IL-6 and IL-8, neutrophil extracellular traps (NETs), and anti-IFN autoantibodies have been identified, scientific understanding of the association is incomplete; systematic and unbiased effort is needed.
We have obtained two-year UNC electronic health record data for approximately 9,000 COVID-19 positive cases (IRB Number: 20-2103) from the Carolina Data Warehouse (CDW) Operations Committee. We are applying machine learning methods to systematically and unbiasedly mine the COVID-19 data to try to identify clinical features that are associated with the severity of COVID-19 infection. We categorized the COVID-19 positive patients in the cohort into four different categories - asymptomatic, mild, severe/critical and death. We will initially employ the tree-based approaches for this dataset. The tree-based approaches are ideal for electronic health records data as those data contain mixed data types including demographics, diagnoses, problem lists, medications, vital signs, and laboratory results. We are particularly interested in clinical features that are predictive of COVID-19 severity. Specifically, we focus on two main questions a) whether sleep apnea is associated with COVID severity; b) whether nutritional deficiencies, such as deficiencies in vitamins B12 and D, are associated with COVID severity. For these analyses, we will use ordinal logistic regression with COVID-19 severity class as the outcome and either sleep apnea (dichotomized as present, absent) or vitamin levels measured in serum as predictors and will adjust for covariates such as age, gender, and BMI when appropriate. We are very excited about this dataset. We are particularly excited about the prospect of interacting with CDW for additional data to support any relationships we discover.
Terms: <2019 novel corona virus><2019 novel coronavirus><2019-nCoV><3-10C><ACE2><AMCF-I><Address><Age><Algorithms><Animals><Area><Area Under Curve><Assay><Autoantibodies><B cell differentiation factor><B cell stimulating factor 2><B-Cell Differentiation Factor><B-Cell Differentiation Factor-2><B-Cell Stimulatory Factor-2><BCDF><BMI><BMI percentile><BMI z-score><BSF-2><BSF2><Basal Transcription Factor><Basal transcription factor genes><Binding><Bio-Informatics><Bioassay><Bioinformatics><Biologic Assays><Biological Assay><Biological Markers><Blood Glucose><Blood Neutrophil><Blood Polymorphonuclear Neutrophil><Blood Serum><Blood Sugar><Body Tissues><Body mass index><COVID disease severity><COVID infected patient><COVID patient><COVID positive patient><COVID severity><COVID-19><COVID-19 disease severity><COVID-19 infected patient><COVID-19 infection><COVID-19 patient><COVID-19 positive><COVID-19 positive patient><COVID-19 positivity><COVID-19 severity><COVID-19 therapy><COVID-19 treatment><COVID-19 virus><COVID19><COVID19 disease severity><COVID19 infection><COVID19 patient><COVID19 positive><COVID19 positive patient><COVID19 positivity><COVID19 severity><COVID19 therapy><COVID19 treatment><COVID19 virus><CV-19><CV19><CXCL8><Cancer cell line><Cancers><Categories><Cell Body><Cell Line><CellLine><Cells><Cessation of life><Chemical Exposure><Clinical><Clinical Trials><CoV-2><CoV2><Consumption><Cyanocobalamin><D-Glucose><DNA Methylation><Data><Data Bases><Data Set><Databases><Dataset><Death><Dextrose><Diagnosis><Disease Management><Disease Outcome><Disorder Management><Drug toxicity><Drug usage><Drugs><Electronic Health Record><Encyclopedias><Exhibits><Exposure to><Expression Signature><Fasting><Fats><Fatty acid glycerol esters><GCP1><GTEx><Gender><Gene Expression><Gene Expression Profile><General Transcription Factor Gene><General Transcription Factors><Generalized Growth><Genes><Genomics><Genotype-Tissue Expression Project><Glucose><Goals><Growth><HPGF><Hepatocyte-Stimulating Factor><High Fat Diet><Human><Humulin R><Hybridoma Growth Factor><IFN><IFN-beta 2><IFNB2><IL-6><IL-8><IL6 Protein><IL8><IL8 gene><IRB><IRBs><Individual><Inflammatory><Institutional Review Boards><Insulin><Interferons><Interleukin-6><K60><Laboratories><Laboratory Scientists><Leptin><Liver><Logistic Regressions><MGI-2><Malignant Neoplasms><Malignant Tumor><Malnutrition><Manuscripts><Marrow Neutrophil><Measures><Medication><Metabolic><Metabolic Diseases><Metabolic Disorder><Metastasis><Metastasize><Metastatic Lesion><Metastatic Mass><Metastatic Neoplasm><Metastatic Tumor><Methodology><Mice><Mice Mammals><Mining><Modeling><Modern Man><Molecular Interaction><Murine><Mus><Muscle><Muscle Tissue><Myeloid Differentiation-Inducing Protein><Neoplasm Metastasis><Neutrophilic Granulocyte><Neutrophilic Leukocyte><Normal Tissue><Normal tissue morphology><Novolin R><Nutritional Deficiency><Ob Gene Product><Ob Protein><Obese Gene Product><Obese Protein><Obesity><Outcome><Over weight><Overweight><Patient Care><Patient Care Delivery><Pharmaceutic Preparations><Pharmaceutical Preparations><Phase><Plasmacytoma Growth Factor><Polymorphonuclear Cell><Polymorphonuclear Leukocytes><Polymorphonuclear Neutrophils><Population><Prevalence><Process><Quetelet index><RNA Seq><RNA sequencing><RNAseq><Receptor Protein><Regular Insulin><SARS corona virus 2><SARS-CoV-2><SARS-CoV-2 disease severity><SARS-CoV-2 infected patient><SARS-CoV-2 infection><SARS-CoV-2 patient><SARS-CoV-2 positive><SARS-CoV-2 positive patient><SARS-CoV-2 positivity><SARS-CoV-2 severity><SARS-CoV-2 therapy><SARS-CoV-2 treatment><SARS-CoV2><SARS-CoV2 infection><SARS-associated corona virus 2><SARS-associated coronavirus 2><SARS-coronavirus-2><SARS-related corona virus 2><SARS-related coronavirus 2><SARSCoV2><SCYB8><Sampling><Secondary Neoplasm><Secondary Tumor><Serum><Severe Acute Respiratory Distress Syndrome CoV 2><Severe Acute Respiratory Distress Syndrome Corona Virus 2><Severe Acute Respiratory Distress Syndrome Coronavirus 2><Severe Acute Respiratory Syndrome CoV 2><Severe Acute Respiratory Syndrome-associated coronavirus 2><Severe Acute Respiratory Syndrome-related coronavirus 2><Severe acute respiratory syndrome associated corona virus 2><Severe acute respiratory syndrome corona virus 2><Severe acute respiratory syndrome coronavirus 2><Severe acute respiratory syndrome coronavirus 2 infection><Severe acute respiratory syndrome coronavirus 2 positive><Severe acute respiratory syndrome coronavirus 2 positivity><Severe acute respiratory syndrome related corona virus 2><Sleep Apnea><Sleep Apnea Syndromes><Sleep Hypopnea><Sleep-Disordered Breathing><Statistical Methods><Strains Cell Lines><Subgroup><TCGA><TSG-1><The Cancer Genome Atlas><Therapeutic><Thesaurismosis><Tissue Growth><Tissues><Transcription Factor Proto-Oncogene><Transcription Regulation><Transcription factor genes><Transcriptional Control><Transcriptional Regulation><Trees><Tumor Tissue><Undernutrition><VIT B12><VIT D><Vitamin B 12><Vitamin B12><Vitamin D><Vitamins><Weight><Weight Gain><Weight Increase><Work><Wuhan coronavirus><adiposity><ages><angiotensin converting enzyme 2><angiotensin converting enzyme II><autoimmune antibody><autoreactive antibody><b-ENAP><base><beta CoV><beta coronavirus><betaCoV><betacoronavirus><bio-markers><biologic marker><biomarker><body weight gain><body weight increase><cancer metastasis><clinical relevance><clinically relevant><cohort><computer based prediction><corona virus disease 2019><coronavirus disease 2019><coronavirus disease 2019 disease severity><coronavirus disease 2019 infected patient><coronavirus disease 2019 infection><coronavirus disease 2019 patient><coronavirus disease 2019 positive><coronavirus disease 2019 positive patient><coronavirus disease 2019 positivity><coronavirus disease 2019 severity><coronavirus disease 2019 therapy><coronavirus disease 2019 treatment><coronavirus disease 2019 virus><coronavirus disease infected patient><coronavirus disease patient><coronavirus disease positive patient><coronavirus disease severity><coronavirus patient><corpulence><cultured cell line><cytokine><data base><data warehouse><demographics><diabetic><diet control><dietary><dietary control><dietary deficiency><drug sensitivity><drug use><drug/agent><effective therapy><effective treatment><electronic health care record><electronic healthcare record><extracellular><fasted><fasts><gene expression pattern><gene expression signature><gene interaction><genome scale><genome-wide><genomewide><hCoV19><hepatic body system><hepatic organ system><high dimensional data><human tissue><impaired glucose tolerance><infected with COVID-19><infected with COVID19><infected with SARS-CoV-2><infected with SARS-CoV2><infected with coronavirus disease 2019><infected with severe acute respiratory syndrome coronavirus 2><interest><interferon beta 2><machine learning method><machine learning methodologies><malignancy><malnourished><metabolic profile><metabolism disorder><multidimensional data><multidimensional datasets><muscular><nCoV2><neoplasm/cancer><neutrophil><new drug treatments><new drugs><new therapeutics><new therapy><next generation therapeutics><novel drug treatments><novel drugs><novel therapeutics><novel therapy><nutrition deficiency><nutrition deficiency disorder><nutritional deficiency disorder><ontogeny><operation><patient infected with COVID><patient infected with COVID-19><patient infected with SARS-CoV-2><patient infected with coronavirus disease><patient infected with coronavirus disease 2019><patient infected with severe acute respiratory syndrome coronavirus 2><patient with COVID><patient with COVID-19><patient with COVID19><patient with SARS-CoV-2><patient with coronavirus disease><patient with coronavirus disease 2019><patient with severe acute respiratory distress syndrome coronavirus 2><prediction model><predictive biomarkers><predictive marker><predictive modeling><predictive molecular biomarker><receptor><response><self reactive antibody><severe acute respiratory syndrome coronavirus 2 disease severity><severe acute respiratory syndrome coronavirus 2 infected patient><severe acute respiratory syndrome coronavirus 2 patient><severe acute respiratory syndrome coronavirus 2 positive patient><severe acute respiratory syndrome coronavirus 2 severity><severe acute respiratory syndrome coronavirus 2 therapy><severe acute respiratory syndrome coronavirus 2 treatment><sleep-related breathing disorder><transcription factor><transcriptional profile><transcriptional signature><transcriptome sequencing><transcriptomics><treat COVID-19><treat COVID19><treat SARS-CoV-2><treat coronavirus disease 2019><treat severe acute respiratory syndrome coronavirus 2><tumor><tumor cell metastasis><tumor specificity><web site><website><wt gain><β CoV><β coronavirus><βCoV>