Document text
Principal Investigator: Zhiyong Lu
Organization: NATIONAL LIBRARY OF MEDICINE
Fiscal Year: 2022
Award: $1,791,061
Funding agency: National Library of Medicine
Over the last decade, the online search for biological information has progressed rapidly and has become an integral part of any scientific discovery process. Today, it is virtually impossible to conduct R&D in biomedicine without relying on the kind of Web resources developed and maintained by the NCBI. Indeed, each day millions of users search for biological information via NCBIs updated online PubMed system. However, finding data relevant to a users information need is not always easy. Improving our understanding of the growing population of Entrez users, their information needs and the way in which they meet these needs opens opportunities to improve information services and information access provided by NCBI. The unfortunate arrival of SARS-CoV-2 and the COVID-19 pandemic has led to unprecedented focused biomedical research and new opportunities to distribute the information learned.
The rapid growth of biomedical literature poses a significant challenge for curation and interpretation. This has become more evident during the COVID-19 pandemic. LitCovid, a literature database of COVID-19 related papers in PubMed, has accumulated over 180,000 articles with millions of accesses. Approximately 10,000 new articles are added to LitCovid every month. A main curation task in LitCovid is topic annotation where an article is assigned with up to eight topics, e.g., Treatment and Diagnosis. The annotated topics have been widely used both in LitCovid (e.g., accounting for 18% of total uses) and downstream studies such as network generation. However, it has been a primary curation bottleneck due to the nature of the task and the rapid literature growth. In response, we developed LITMC-BERT, a transformer-based multi-label classification method in biomedical literature. It uses a shared transformer backbone for all the labels while also captures label-specific features and the correlations between label pairs. We compare LITMC-BERT with three baseline models on two datasets. Its micro-F1 and instance-based F1 are 5% and 4% higher than the current best results, respectively, and only requires 18% of the inference time than the Binary BERT baseline.
In addition, we organized the BioCreative LitCovid track to call for a community effort to tackle automated topic annotation for COVID-19 literature. The BioCreative LitCovid dataset consisting of over 30,000 articles with manually reviewed topics was created for training and testing. It is one of the largest multi-label classification datasets in biomedical scientific literature. Nineteen teams worldwide participated and made 80 submissions in total. Most teams used hybrid systems based on transformers. The highest performing submissions achieved 0.8875, 0.9181, and 0.9394 for macro F1-score, micro F1-score, and instance-based F1-score, respectively. Notably, these scores are substantially higher (e.g., 12%, higher for macro F1-score) than the corresponding scores of the state-of-art multi-label classification method.
In 2022, we also benchmarked five DL models, Convolutional Neural Network, BioSentVec, BioBERT, BlueBERT, and ClinicalBERT, for the task of semantic textual similarity (STS). We evaluated a random forest model as an additional baseline. For each model, we repeated the experiment 10 times, using the official training and testing sets. We reported 95% CI of the Wilcoxon rank-sum test on the average Pearson correlation (official evaluation metric) and running time. We further evaluated Spearman correlation, R, and mean squared error as additional measures. Using only the official training set, all models obtained highly effective results. BioSentVec and BioBERT achieved the highest average Pearson correlations (0.8497 and 0.8481, respectively).
Finally, we made use of simple natural language processing programs and robust statistical tests in a collaborative project with researchers at NHGRI studying the evolving use of ancestry, ethnicity, and race in genetics research. Our computational method allowed them to analyze tens of thousands of pages easily and find associations between words. Similarly, in another collaboration with NCI researchers, we applied our machine learning research to classify literature and extract data at the intersection of three fields: liver cancer, health disparities, and epidemiology.
Terms: <2019 novel corona virus><2019 novel coronavirus><2019-nCoV><Accounting><Benchmarking><Best Practice Analysis><Biological><Biomedical Research><COVID crisis><COVID epidemic><COVID pandemic><COVID-19><COVID-19 crisis><COVID-19 epidemic><COVID-19 global health crisis><COVID-19 global pandemic><COVID-19 health crisis><COVID-19 pandemic><COVID-19 public health crisis><COVID-19 virus><COVID19><COVID19 crisis><COVID19 epidemic><COVID19 global health crisis><COVID19 global pandemic><COVID19 health crisis><COVID19 pandemic><COVID19 public health crisis><COVID19 virus><CV-19><CV19><Classification><CoV-2><CoV2><Collaborations><Communities><Computing Methodologies><ConvNet><DNA Molecular Biology><Data><Data Bases><Data Set><Databases><Dataset><Development and Research><Diagnosis><Epidemiology><Ethnic Origin><Ethnicity><Evaluation><Generalized Growth><Generations><Genetic Research><Goals><Growth><Hepatic Cancer><Hybrids><Information Services><Internet><Investigators><Label><Link><Literature><Machine Learning><Malignant neoplasm of liver><Manuals><Measures><Methods><Modeling><Molecular Biology><NHGRI><National Center for Human Genome Research><National Human Genome Research Institute><Natural Language Processing><Nature><Paper><Population><Process><PubMed><R & D><R&D><Race><Racial Group><Racial Stocks><Rank-Sum Tests><Reporting><Research><Research Personnel><Researchers><Running><SARS corona virus 2><SARS-CO-V2><SARS-COVID-2><SARS-CoV-2><SARS-CoV-2 epidemic><SARS-CoV-2 global health crisis><SARS-CoV-2 global pandemic><SARS-CoV-2 pandemic><SARS-CoV2><SARS-CoV2 epidemic><SARS-CoV2 pandemic><SARS-associated corona virus 2><SARS-associated coronavirus 2><SARS-coronavirus-2><SARS-coronavirus-2 epidemic><SARS-coronavirus-2 pandemic><SARS-related corona virus 2><SARS-related coronavirus 2><SARSCoV2><Semantics><Severe Acute Respiratory Coronavirus 2><Severe Acute Respiratory Distress Syndrome CoV 2><Severe Acute Respiratory Distress Syndrome Corona Virus 2><Severe Acute Respiratory Distress Syndrome Coronavirus 2><Severe Acute Respiratory Syndrome CoV 2><Severe Acute Respiratory Syndrome CoV 2 epidemic><Severe Acute Respiratory Syndrome CoV 2 pandemic><Severe Acute Respiratory Syndrome-associated coronavirus 2><Severe Acute Respiratory Syndrome-related coronavirus 2><Severe acute respiratory syndrome associated corona virus 2><Severe acute respiratory syndrome corona virus 2><Severe acute respiratory syndrome coronavirus 2><Severe acute respiratory syndrome coronavirus 2 epidemic><Severe acute respiratory syndrome coronavirus 2 pandemic><Severe acute respiratory syndrome related corona virus 2><Spinal Column><Spine><System><Systematics><Testing><Time><Tissue Growth><Training><Update><Vertebral column><WWW><Wuhan coronavirus><backbone><base><biologic><cancer disparity><cancer health disparity><cancer-related health disparity><computational methodology><computational methods><computer based method><computer methods><computing method><convolutional network><convolutional neural nets><convolutional neural network><corona virus disease 2019><corona virus disease 2019 epidemic><corona virus disease 2019 pandemic><coronavirus disease 2019><coronavirus disease 2019 crisis><coronavirus disease 2019 epidemic><coronavirus disease 2019 global health crisis><coronavirus disease 2019 global pandemic><coronavirus disease 2019 health crisis><coronavirus disease 2019 pandemic><coronavirus disease 2019 public health crisis><coronavirus disease 2019 virus><coronavirus disease crisis><coronavirus disease epidemic><coronavirus disease pandemic><coronavirus disease-19><coronavirus disease-19 global pandemic><coronavirus disease-19 pandemic><coronavirus disease-19 virus><coronavirus infectious disease-19><data base><disparity in cancer><epidemiologic><epidemiological><experiment><experimental research><experimental study><hCoV19><improved><internet resource><liver cancer><liver malignancy><machine learned><malignant liver tumor><nCoV2><natural language understanding><on-line compendium><on-line resource><online compendium><online resource><ontogeny><programs><random forest><rapid growth><research and development><response><severe acute respiratory syndrome coronavirus 2 global health crisis><severe acute respiratory syndrome coronavirus 2 global pandemic><time use><virtual><web><web resource><web services><web-based resource><web-based service><world wide web>