Integrative transcriptomics to uncover functional elements and disease-associated variants in RNA

NIH Pandemic-Era Grants

Pandemic Era Grants

2023

Document text

Principal Investigator: David Anthony Hendrix
Organization: OREGON STATE UNIVERSITY
Fiscal Year: 2023
Award: $305,586
Funding agency: National Institute of General Medical Sciences

PROJECT SUMMARY:
There is a need for integrative data analyses that anchor transcriptomic research in contexts predictive of
human health, as illustrated by growing awareness of disease-associated synonymous transcript variants
and RNA biotechnologies such as mRNA vaccines. To help uncover sequence features that are important
for RNA regulation, we present context-dependent models of translational efficiency, a key metric of
transcript function. We show that position-dependent codon usage bias (PDCUB) identifies start codons
among AUGs more consistently than the Kozak sequence, while high-PDCUB transcripts are enriched for
medically important genes tied to human development and neural function. Attention-based transformer
networks and interpretation techniques will independently predict translational efficiency in human
transcripts, with comparison to ribosome profiling and RNA abundance data in multiple human cell lines,
to characterize how PDCUB and other sequence features guide translational efficiency across health-
critical contexts. Transfection assays validate the roles of predicted sequence features.
 Beyond sequence, higher-order structures also drive RNA function and stability, including
translational regulation and interactions with microRNAs and RNA-binding proteins (RBPs). A new RNA
structural alignment method and associated clustering will uncover structural domains and group them
by mutual similarity to find common structural motifs that impact RNA structure-function relationships,
improving our understanding of the role of transcript structure in pathogenesis. Evaluation will consist of
clustering RNA families in our previously built RNA structure meta-database, bpRNA-1m, with identified
structural domains analyzed in the context of ribosome profiling data to characterize the role of these
domains in regulating translation. Meanwhile, clustering structures according to RNA-protein crosslinking
data will let us identify motifs involved in the binding of RBPs.
 Finally, a comprehensive transcriptome browser and meta-database will integrate transcriptomic
data for known and new transcript-level features, including those described above. Easy to access and
use, this resource will enable scientific and medical researchers to find and define RNA sequence features
and structural motifs. By cohesively cataloging the complex facets of transcript-level interactions, along
with sequence and structural features relevant for transcript regulation, our transcriptome browser will
help researchers visualize ribosomal occupancy, examine RNA structures, microRNA and RBP binding,
catalog splice variants, and understand the sequence features that drive transcript interactions. Allelic
variants mapped to RNA transcript positions will be combined our annotations, along with feature-based
machine learning predictions incorporated into the browser, to assist researchers in generating first-pass
predictions of transcript variants and interpreting their outcomes in the context of human health.

Terms: <Algorithms><Assay><Attention><Awareness><Binding><Binding Proteins><Bioassay><Biologic Assays><Biological><Biological Assay><Biotech><Biotechnology><CLIP-Seq><Cataloging><Catalogs><Code><Coding System><Codon><Codon Nucleotides><Complex><Custom><Data><Data Analyses><Data Analysis><Data Bases><Data Set><Databases><Development><Disease><Disorder><Elements><Evaluation><Family><Foundations><Gene variant><Genes><Genome><Genomics><HITS-CLIP><Health><High-throughput sequencing of CLIP cDNA library><Higher Order Chromatin Folding><Higher Order Chromatin Structure><Higher Order Structure><Human><Human Cell Line><Human Development><In Vitro><Infrastructure><Initiation Codon><Initiation Factors><Initiator Codon><Investigation><Investigators><JTK14><Life><Ligand Binding Protein><Ligand Binding Protein Gene><Location><Machine Learning><Maps><Measures><Medical><Methods><Micro RNA><MicroRNAs><Modeling><Modern Man><Molecular Interaction><Neurophysiology - biologic function><Non-Polyadenylated RNA><Outcome><Output><Pathogenesis><Pattern><Peptide Initiation Factors><Position><Positioning Attribute><Process><Protein Binding><Proteins><RNA><RNA Binding><RNA Gene Products><RNA Seq><RNA Sequences><RNA Splicing><RNA bound><RNA sequencing><RNA vaccine><RNA-Binding Proteins><RNA-based vaccine><RNAseq><Regulation><Research><Research Personnel><Research Resources><Researchers><Resources><Ribo-seq><Ribonucleic Acid><Ribosomal RNA><Ribosomes><Role><Site><Splicing><Standardization><Start Codon><Structure><Structure-Activity Relationship><TIE gene><TIE1><Techniques><Training><Transcript><Transfection><Translation Initiation Factor><Translational Initiation Factor><Translational Regulation><Translations><Update><Variant><Variation><Visualization><Work><allele variant><allelic variant><biologic><bound protein><catalog><cell type><chemical structure function><crosslinking and immunoprecipitation sequencing><customs><data base><data interpretation><developmental><genetic variant><genomic variant><global gene expression><global transcription profile><improved><insight><intermolecular interaction><mRNA Stability><mRNA Translation><mRNA vaccine><mRNA-based vaccine><machine based learning><machine learning based method><machine learning based prediction model><machine learning based predictive model><machine learning method><machine learning methodologies><machine learning prediction><machine learning prediction model><miRNA><miRNAs><neural function><novel><protein crosslink><rRNA><ribosome footprint profiling><ribosome profiling><social role><structure function relationship><transcriptome><transcriptome sequencing><transcriptomic sequencing><transcriptomics><translation><translational model><user-friendly><web-accessible>