Methods for single-cell CRISPR screens and multiomic data: constructing powerful well-calibrated tests, circumventing unmeasured confounding, and accounting for denoising and imputation

NIH Pandemic-Era Grants

Pandemic Era Grants

2024

Document text

Principal Investigator: KATHRYN M ROEDER
Organization: CARNEGIE-MELLON UNIVERSITY
Fiscal Year: 2024
Award: $592,551
Funding agency: National Institute of Mental Health

Project summary/abstract
In this application, we request continuation of MH123184, which aimed to understand how
genetic variation alters transcription in specific cells and thereby produces psychopathology. Our research
developed statistical methods to integrate single cell and tissue-level transcriptomic data. We targeted
methods to identify gene communities, defined in terms of cell type and spatiotemporal window, to
understand how genes act in concert to confer risk for psychopathology. We also took advantage of an
exciting new avenue of research to approach these challenges, namely CRISPR screening. This innovation
has emerged as a powerful tool to characterize the effects of genetic perturbations on the entire
transcriptome at a single-cell level. Here we propose research covering three related themes, all of which
capitalize on CRISPR advancements: (1) develop powerful and well-calibrated tests for the effect of CRISPR
perturbations on gene expression by inferring latent factors; (2) develop methods for removing the effect of
unmeasured confounders in high throughput screens; and (3) develop methods for imputation and
denoising for multiomic data that facilitate downstream testing of omic readouts. Each of these aims is
motivated by pressing needs in the field. First, due to small samples and the sparsity of the response
variable, it is essential that we enhance the power and interpretability of CRISPR tests by accounting for co-
regulation and convergent function of genes. Aim 1 achieves this purpose by estimating latent factors that
represent co-regulated genes and by inferring a similarity matrix among gene perturbations. As CRISPR
screens advance to more biologically complex settings, such as model organisms, unmeasured confounders
will play a more important role, and new methods are needed to control for these effects. Aim 2 develops
two approaches to this challenge: an innovative use of negative control variables, as motivated by the
causal literature, and key advances to the classic surrogate variable analysis method. For the field to move
toward efficient use of multiomic data, data derived from multiple sources will be required. These
resources will invariably have missing data. Methods to account for imputation of missing data are needed.
Tools developed for variational autoencoders show great promise; however, as described in Aim 3, they
need to be paired with semiparametric inference tools to ensure robust and well calibrated downstream
analysis. By applying what we learn from these three aims to available resources, most from distributed
resources and some from our collaborations, we expect to shed more light on the neurobiological
mechanisms of mental illness. We are well positioned to move between theory and data because
we have a diverse team of investigators lead by the PI (Roeder), who has decades of experience
in statistical genomic field and co-investigators Wasserman and Lei, who are experts in theory and
methods for high dimensional and causal inference. The proposed research fits Goal 1 of NIMH’s
Strategic Plan, advancing basic science of brain, genomics, and behavior to understand mental
illnesses.

Terms: <Acceleration><Accounting><Affect><Animal Model><Animal Models and Related Studies><Atlases><Basic Research><Basic Science><Behavior><Body Tissues><Brain><Brain Nervous System><CRISPR><CRISPR editing screen><CRISPR screen><CRISPR-based screen><CRISPR/Cas system><CRISPR/Cas9 screen><Calibration><Causality><Cell Body><Cells><Chromatin><Clustered Regularly Interspaced Short Palindromic Repeats><Collaborations><Communities><Complex><Computer software><Coupled><Data><Data Set><Development><Encephalon><Ensure><Etiology><Functional RNA><Gene Expression><Gene Proteins><Gene Transcription><Genes><Genetic><Genetic Diversity><Genetic Risk><Genetic Transcription><Genetic Variation><Genetic predisposing factor><Genomics><Goals><Health><High Throughput Assay><Human><In Vitro><Individual><Investigation><Investigators><Joints><Lead><Learning><Literature><Measurement><Mental disorders><Mental health disorders><Methods><Modeling><Modern Man><Molecular><Multiomic Data><NIMH><National Institute of Mental Health><Nerve Cells><Nerve Unit><Neural Cell><Neurocyte><Neurons><Non-Coding><Non-Coding RNA><Non-translated RNA><Noncoding RNA><Nontranslated RNA><Nucleic Acid Regulator Regions><Nucleic Acid Regulatory Sequences><Outcome><Pb element><Peptides><Personal Satisfaction><Play><Population><Position><Positioning Attribute><Procedures><Protein Gene Products><Proteins><Proteomics><Psychiatric Disease><Psychiatric Disorder><Psychopathology><QTL><Quantitative Trait Loci><RNA Expression><Regulatory Regions><Research><Research Personnel><Research Resources><Researchers><Resources><Risk><Risk-associated variant><Role><Sampling><Software><Source><Statistical Methods><Strategic Planning><Structure><Supporting Cell><Testing><Tissues><Transcription><Untranslated RNA><Variant><Variation><abnormal psychology><analytical method><autoencoder><autoencoding neural network><causation><cell type><clustered regularly interspaced short palindromic repeats screen><data imputation><data integration><de-noising><deep learning><deep learning method><deep learning strategy><denoising><design><designing><developmental><differential expression><differentially expressed><disease causation><experience><experiment><experimental research><experimental study><experiments><gene function><genetic regulatory element><genetic risk factor><genome scale><genome-wide><genomewide><global gene expression><global transcription profile><heavy metal Pb><heavy metal lead><high dimensionality><high throughput screening><improved><imputation method><in vivo><inherited factor><innovate><innovation><innovative><knock-out animal><knockout animal><large scale data><large scale data sets><large scale datasets><machine learning based method><machine learning method><machine learning methodologies><machine statistical learning><mental illness><model of animal><model organism><mosaic><multiomics><multiple omic data><multiple omics><neurobiological mechanism><neuronal><noncoding><panomics><psychiatric illness><psychological disorder><response><risk allele><risk gene><risk genotype><risk loci><risk locus><risk variant><scRNA-seq><semiparametric><single cell RNA-seq><single cell RNAseq><single cell analysis><single cell expression profiling><single cell technology><single cell transcriptomic profiling><single-cell RNA sequencing><social role><spatiotemporal><statistic methods><statistical and machine learning><theories><tool><transcriptional differences><transcriptome><transcriptomics><user friendly computer software><user friendly software><well-being><wellbeing>