Document text
Principal Investigator: JAMES C MULLIKIN
Organization: NATIONAL HUMAN GENOME RESEARCH INSTITUTE
Fiscal Year: 2022
Award: $10,295,725
Funding agency: National Human Genome Research Institute
Over the last year, NISC operated the following suite of production sequencing machines: 1 PacBio Sequel 2, 1 NovaSeq 6000, 1 Oxford Nanopore GridION, 1 Oxford Nanopore PromethION, and 3 MiSeqs. Using these platforms, we have generated over 3,259 billion reads in the past year. Though we remain consistently at a level of a mid-scale genome sequencing center, we have maintained advantageous economies of scale while remaining relatively agile. The NovaSeq 6000 allows NISC to effectively meet the rising interest in studying whole genome sequence datasets, with over 1,634 human genomes sequenced and analyzed in the past year. Using our long-read sequencing technologies, we contributed to the Science Journal publication of the complete, telomere-to-telomere sequence of a human genome. (Nurk, Koren et al. 2022)
The adoption of many new sequencing protocols in production created the commensurate need for dramatic changes to sample tracking, flow control and primary analysis pipelines, as well as project management and cost accounting. Ongoing rapid design, development, and implementation of new Laboratory Information Management System (LIMS) features by a dedicated NISC team allows the system to evolve quickly to adapt to a continuous flow of changes in sequencing technologies. A combination of talented IT staff and bioinformaticians have met the challenges of extremely large and complex data sets by implementing and continuously adapting pipeline programs to support rapidly evolving software associated with each of the sequencing platforms. Beyond primary analysis that results in DNA basecalls and quality scores, NISC has worked closely with members of other NHGRI research groups to implement and support high-throughput production of biologically relevant secondary analysis. One shining example of these efforts is the production scale processing of Whole Genome Sequencing (WGS) data for all our clients, using a newly implemented GPU accelerated GATK4 best practices pipeline. With these new sequence alignment and genetic variant calling systems in place, we can analyze WGS datasets at a rate of one per hour, matching our maximum throughput of the NovaSeq 6000.
Publications for fiscal-year 2022 span a wide range of projects, and are summarized as follows:
1) WES projects (n = 4) (Drazer, Homan et al. 2022; Li, Yang et al. 2022; Pitsava, Feldkamp et al. 2022; Rudd, Hansen et al. 2022)
2) Whole Genome Sequencing, Assembly and Analysis (n = 2) Altemose, Logsdon et al. 2022; Nurk, Koren et al. 2022)
3) Microbiome study (n = 3) (Jo, Harkins et al. 2021; Kashaf, Proctor et al. 2022; Sim, Kashaf et al. 2022)
The NISC team worked with two PIs at NIH related to COVID-19 and SARS-CoV-2. In collaboration with Dr. Cliff Lane at NIAID we have generated 2,182 SARS-CoV-2 genomes from patients in Clinical Trials. In collaboration with Dr. Lothar Hennighausen at NIDDK, we have generated 568 RNA-seq datasets from patients in COVID-19 studies.
In the foreseeable future, NISC is well positioned to provide next-gen sequence data for a multitude of investigators across NIH. We also expect increasing access to sequencing by the NIH Clinical Center with our CLIA exome test and continuing our sequencing support for Intramural NHGRI investigators for their most promising projects. Our focus is to increase operational efficiencies of the next-gen pipeline, refine existing protocols, implement additional protocols as new sample/experimental types are requested from researchers and continue to expand the value-added data analysis packages available. With our new PromethION fully integrated into our production pipeline along side HiFi PacBio long-read sequences, the generation of super high-quality and contiguous genome assemblies are available to NIH intramural researchers. In summary, we will continue to monitor developments in the rapidly evolving sequencing and informatics technologies, implementing those we deem most appropriate for our collaborating investigators.
Terms: <2019 novel corona virus><2019 novel coronavirus><2019-nCoV><Accounting><Adoption><Basic Research><Basic Science><Biology><COVID-19><COVID-19 genome><COVID-19 virus><COVID-19 virus genome><COVID19><COVID19 genome><COVID19 virus><COVID19 virus genome><CV-19><CV19><ChIP Sequencing><ChIP-seq><Client><Clinical Sciences><Clinical Trials><CoV-2><CoV2><Collaborations><Computer software><Custom><DNA><DNA seq><DNA sequencing><DNAseq><Data><Data Analyses><Data Analysis><Data Set><Dataset><Deoxyribonucleic Acid><Development><Future><Gene variant><Generations><Genome><Goals><Hour><Human Genome><Informatics><Infrastructure><Intramural Program><Intramural Research Program><Investigators><Journals><Laboratories><Magazine><Management Information Systems><Medicine><Methods><Microbiomics><Monitor><NHGRI><NIAID><NIDDK><NIH><National Center for Human Genome Research><National Human Genome Research Institute><National Institute of Allergy and Infectious Disease><National Institute of Diabetes and Digestive and Kidney Diseases><National Institutes of Health><Patients><Position><Positioning Attribute><Production><Protocol><Protocols documentation><Publications><RNA Seq><RNA sequencing><RNAseq><Research><Research Personnel><Researchers><Role><SARS corona virus 2><SARS-CO-V2><SARS-COVID-2><SARS-CoV-2><SARS-CoV-2 genome><SARS-CoV2><SARS-CoV2 genome><SARS-associated corona virus 2><SARS-associated coronavirus 2><SARS-coronavirus-2><SARS-related corona virus 2><SARS-related coronavirus 2><SARSCoV2><SEQ-AN><Sampling><Science><Scientific Publication><Sequence Alignment><Sequence Analyses><Sequence Analysis><Severe Acute Respiratory Coronavirus 2><Severe Acute Respiratory Distress Syndrome CoV 2><Severe Acute Respiratory Distress Syndrome Corona Virus 2><Severe Acute Respiratory Distress Syndrome Coronavirus 2><Severe Acute Respiratory Syndrome CoV 2><Severe Acute Respiratory Syndrome-associated coronavirus 2><Severe Acute Respiratory Syndrome-related coronavirus 2><Severe acute respiratory syndrome associated corona virus 2><Severe acute respiratory syndrome corona virus 2><Severe acute respiratory syndrome coronavirus 2><Severe acute respiratory syndrome related corona virus 2><Side><Software><System><Talents><Technology><Testing><United States National Institutes of Health><Work><Wuhan coronavirus><Yang><allele variant><allelic variant><analysis pipeline><chromatin immunoprecipitation-sequencing><clinical center><comparative><complex data><corona virus disease 2019><coronavirus disease 2019><coronavirus disease 2019 genome><coronavirus disease 2019 virus><coronavirus disease 2019 virus genome><coronavirus disease-19><coronavirus disease-19 virus><coronavirus infectious disease-19><cost><data interpretation><design><designing><developmental><entire genome><exome><exome sequencing><exome-seq><exomes><full genome><genetic variant><genome scale><genome sequencing><genome-wide><genomewide><genomic data><genomic data-set><genomic dataset><genomic variant><hCoV19><human whole genome><interest><meetings><member><microbiome research><microbiome science><microbiome studies><nCoV2><nano pore><nanopore><novel><programs><secondary analysis><sequencing alignment><sequencing platform><severe acute respiratory syndrome coronavirus 2 genome><social role><telomere><transcriptome sequencing><whole genome>