Document text
Principal Investigator: CATHY H. WU
Organization: UNIVERSITY OF DELAWARE
Fiscal Year: 2024
Award: $433,379
Funding agency: National Institute of General Medical Sciences
Protein Knowledge Networks and Semantic Computing for Disease Discovery
The growing volume and breadth of information from the scientific literature and biomedical databases
pose challenges to the research community to exploit the content for discovery. This MIRA grant
application will advance our knowledge mining and semantic computing system to accelerate data-driven
discovery for understanding of gene-disease-drug relationships. We have employed natural language
processing and machine learning approaches in a generalizable framework for bioentity and relation
extraction from large-scale text. Our Protein Ontology supports protein-centric semantic integration of
biomedical data for both human understanding and computational reasoning. We have also developed a
resource to support functional interpretation and analysis of protein post-translational modifications
(PTMs) across modification types and organisms. Building on our computational algorithms,
bioinformatics infrastructure and community interactions, we will further develop literature mining tools to
support automated information extraction across the bibliome and open linked data models for semantic
integration of biomedical data from heterogeneous resources. Our text mining tools will be trained for
different use cases using deep learning methods. We will develop RDF (Resource Description
Framework) semantic models in an increasingly computable, inferable and explainable knowledge
system to assist in hypothesis generation. We will present evidence in the form of textual artifacts and
semantic models to ensure unbiased analysis and interpretation of results to promote rigorous and
reproducible research. We will develop scientific case studies to drive the system development.
Examples include PTM disease variant and enrichment analyses for drug target identification, genotype-
phenotype knowledge mining for Alzheimer's Disease understanding, and gene-disease-drug knowledge
network construction for COVID-19 drug repurposing. To foster community engagement, we will host
workshops and hackathons to address critical fundamental research questions and emerging disease
scenarios. We have fully adopted the FAIR (Findable, Accessible, Interoperable, Reusable) principles for
resource sharing. All data, tools and research results will be broadly disseminated from the project
website, accessible programmatically via RESTful API, queryable via SPARQL endpoints, and
dockerized for community code reuse. The successful completion of this research will thus support
scalable, integrative and collaborative knowledge discovery to accelerate disease understanding and
drug target discovery.
Terms: <AD dementia><Acceleration><Address><Adopted><Alzheimer Type Dementia><Alzheimer disease dementia><Alzheimer sclerosis><Alzheimer syndrome><Alzheimer's><Alzheimer's Disease><Alzheimers Dementia><Applications Grants><Artifacts><Biomedical Research><COVID-19><CV-19><Case Study><Code><Coding System><Communities><Computational algorithm><Computer Systems><Coronavirus Infectious Disease 2019><Data><Data Bases><Databases><Disease><Disorder><Drug Targeting><Drugs><Educational workshop><Ensure><FAIR data><FAIR guiding principles><FAIR principles><Findable, Accessible, Interoperable and Re-usable><Findable, Accessible, Interoperable, and Reusable><Fostering><Generalized Growth><Generations><Genes><Genotype><Grant Proposals><Growth><Health><Human><Information Retrieval><Information extraction><Knowledge><Knowledge Discovery><Link><Literature><Machine Learning><Medication><Mining><Modeling><Modern Man><Modification><Morphologic artifacts><Natural Language Processing><Ontology><Organism><Pharmaceutical Preparations><Phenotype><Post-Translational Modification Protein/Amino Acid Biochemistry><Post-Translational Modifications><Post-Translational Protein Modification><Post-Translational Protein Processing><Posttranslational Modifications><Posttranslational Protein Processing><Primary Senile Degenerative Dementia><Protein Analysis><Protein Modification><Proteins><Reproducibility><Research><Research Resources><Resource Description Framework><Resource Sharing><Resources><Scientist><Semantics><System><Systems Development><Text><Tissue Growth><Training><Variant><Variation><Workshop><bio-informatics infrastructure><bioinformatics infrastructure><case report><community engagement><computational reasoning><computational thinking><computer algorithm><computing system><coronavirus disease 2019><coronavirus disease-19><coronavirus infectious disease-19><data base><data modeling><deep learning><deep learning method><deep learning strategy><discovery mining><drug repositioning><drug repurposing><drug/agent><engagement with communities><fundamental research><hack day><hackathon><hackfest><human disease><improved><knowledge network><literature mining><literature searching><living system><machine based learning><model of data><model the data><modeling of the data><natural language understanding><ontogeny><pandemic><pandemic disease><primary degenerative dementia><repurposing agent><repurposing medication><senile dementia of the Alzheimer type><text mining><text searching><tool><web site><website>