Document text
Principal Investigator: James Thomas
Organization: NATIONAL HUMAN GENOME RESEARCH INSTITUTE
Fiscal Year: 2023
Award: $9,722,540
Funding agency: National Human Genome Research Institute
Over the last year, NISC operated the following suite of production sequencing machines: 2 PacBio Sequel II, 1 Illumina NovaSeq 6000, 1 Oxford Nanopore GridION, 1 Oxford Nanopore P24 PromethION, and 4 Illumina MiSeqs. Using these platforms, we have generated over 3,222 billion reads in the past year. Though we remain consistently at a level of a mid-scale genome sequencing center, we have maintained advantageous economies of scale while remaining relatively agile. The NovaSeq 6000 allows NISC to effectively meet the rising interest in studying whole genome sequence datasets, with over 2,032 human genomes sequenced and analyzed in the past year. Using our long-read sequencing technologies, we continue to contribute to projects aimed at generating complete (telomere-to -telomere, T2T) genome assemblies of humans and other animals.
The adoption of many new sequencing protocols in production created the commensurate need for dramatic changes to sample tracking, flow control and primary analysis pipelines, as well as project management and cost accounting. Ongoing rapid design, development, and implementation of new Laboratory Information Management System (LIMS) features by a dedicated NISC team allows the system to evolve quickly to adapt to a continuous flow of changes in sequencing technologies. A combination of talented IT staff and bioinformaticians have met the challenges of extremely large and complex data sets by implementing and continuously adapting pipeline programs to support rapidly evolving software associated with each of the sequencing platforms. Beyond primary analysis that results in DNA basecalls and quality scores, NISC has worked closely with members of other NHGRI research groups to implement and support high-throughput production of biologically relevant secondary analysis. One shining example of these efforts is the production scale processing of Whole Genome Sequencing (WGS) data for all our clients, using a newly implemented GPU accelerated GATK4 best practices pipeline. With these new sequence alignment and genetic variant calling systems in place, we can analyze WGS datasets at a rate of one per hour, matching our maximum throughput of the NovaSeq 6000.
Publications for fiscal-year 2023 span a wide range of projects, and are summarized as follows:
1) Exome and targeted sequencing projects (n = 3) (Loftus et al. 2023; Sok et al. 2023; Simpson et al. 2023)
2) Whole genome sequencing, assembly and analysis (n = 2) (Tisza et al. 2023; Weller et al. 2023)
3) Microbiome study (n = 1) (Saheb et al. 2023)
4) scRNA-seq study (n = 1)(Gordon-Lipkin et al. 2023)
The NISC team worked with two PIs at NIH related to COVID-19 and SARS-CoV-2. In collaboration with Dr. Cliff Lane at NIAID we have generated 1,452 SARS-CoV-2 genomes from patients in Clinical Trials. In collaboration with Dr. Lothar Hennighausen at NIDDK, we have generated 64 RNA-seq datasets from patients in COVID-19 studies.
In the foreseeable future, NISC is well positioned to provide next-gen sequence data for a multitude of investigators across NIH. We also expect increasing access to sequencing by the NIH Clinical Center with our CLIA exome test and continuing our sequencing support for Intramural NHGRI investigators for their most promising projects. Our focus is to increase operational efficiencies of the next-gen pipeline, refine existing protocols, implement additional protocols as new sample/experimental types are requested from researchers and continue to expand the value-added data analysis packages available. With our new PromethION fully integrated into our production pipeline along side HiFi PacBio long-read sequences, super high-quality and contiguous genome assemblies are available to NIH intramural researchers. In summary, we will continue to monitor developments in the rapidly evolving sequencing and informatics technologies, implementing those we deem most appropriate for our collaborating investigators.
Terms: <2019 novel corona virus><2019 novel coronavirus><2019-nCoV><Acceleration><Accounting><Adoption><Animals><Basic Research><Basic Science><Biology><COVID-19><COVID-19 genome><COVID-19 virus><COVID-19 virus genome><COVID19><COVID19 genome><COVID19 virus><COVID19 virus genome><CV-19><CV19><Cell Body><Cells><ChIP Sequencing><ChIP-seq><ChIPseq><Client><Clinical Sciences><Clinical Trials><CoV-2><CoV2><Collaborations><Computer software><DNA><DNA seq><DNA sequencing><DNAseq><Data><Data Analyses><Data Analysis><Data Set><Dedications><Deoxyribonucleic Acid><Development><Future><Gene variant><Genome><Goals><Hour><Human><Human Genome><Informatics><Infrastructure><Intramural Program><Intramural Research Program><Investigators><Laboratories><Management Information Systems><Medicine><Methods><Microbiomics><Modern Man><Monitor><NHGRI><NIAID><NIDDK><NIH><National Center for Human Genome Research><National Human Genome Research Institute><National Institute of Allergy and Infectious Disease><National Institute of Diabetes and Digestive and Kidney Diseases><National Institutes of Health><Patients><Position><Positioning Attribute><Production><Protocol><Protocols documentation><Publications><RNA Seq><RNA sequencing><RNAseq><Research><Research Personnel><Researchers><Role><SARS corona virus 2><SARS-CO-V2><SARS-COVID-2><SARS-CoV-2><SARS-CoV-2 genome><SARS-CoV2><SARS-CoV2 genome><SARS-associated corona virus 2><SARS-associated coronavirus 2><SARS-coronavirus-2><SARS-related corona virus 2><SARS-related coronavirus 2><SARSCoV2><Sampling><Scientific Publication><Sequence Alignment><Severe Acute Respiratory Coronavirus 2><Severe Acute Respiratory Distress Syndrome CoV 2><Severe Acute Respiratory Distress Syndrome Corona Virus 2><Severe Acute Respiratory Distress Syndrome Coronavirus 2><Severe Acute Respiratory Syndrome CoV 2><Severe Acute Respiratory Syndrome-associated coronavirus 2><Severe Acute Respiratory Syndrome-related coronavirus 2><Severe acute respiratory syndrome associated corona virus 2><Severe acute respiratory syndrome coronavirus 2><Severe acute respiratory syndrome related corona virus 2><Side><Software><System><Talents><Technology><Testing><United States National Institutes of Health><Work><Wuhan coronavirus><allele variant><allelic variant><analysis pipeline><chromatin immunoprecipitation-sequencing><clinical center><comparative><complex data><corona virus disease 2019><coronavirus disease 2019><coronavirus disease 2019 genome><coronavirus disease 2019 virus><coronavirus disease 2019 virus genome><coronavirus disease-19><coronavirus disease-19 virus><coronavirus infectious disease-19><cost><data interpretation><design><designing><developmental><entire genome><exome><exome sequencing><exome-seq><exomes><full genome><genetic variant><genome scale><genome sequencing><genome-wide><genomewide><genomic variant><hCoV19><human whole genome><interest><meeting><meetings><member><microbiome research><microbiome science><microbiome studies><nCoV2><nano pore><nanopore><novel><programs><scRNA-seq><secondary analysis><sequencing alignment><sequencing platform><severe acute respiratory syndrome coronavirus 2 genome><single cell RNA-seq><single cell RNAseq><single cell expression profiling><single cell transcriptomic profiling><single-cell RNA sequencing><social role><targeted sequencing><telomere><transcriptome sequencing><transcriptomic sequencing><whole genome>