Document text
FDA-CBER-2021-5683-1072366
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 2TABLE OF CONTENTS
LIST OF TABLES ................................ ................................ ................................ ..................... 2
LIST OF FIGURES ................................ ................................ ................................ ................... 2
1. BACKGROUND ................................ ................................ ................................ ................... 3
1.1. Targeted NGS for SARS -CoV -2................................ ................................ ............... 3
1.2.Objective ................................ ................................ ................................ ................... 3
2. MATERIAL S AND MET HODS ................................ ................................ ........................... 3
2.1. Nucleic Acid Extraction ................................ ................................ ............................ 3
2.2. I on Torrent................................ ................................ ................................ ................. 4
2.2.1. I on Torrent Library Preparati on and Sequencing ................................ ......... 4
2.2.2. I on Torrent Bioinformatics Workflow ................................ .......................... 4
2.3. I llumina ................................ ................................ ................................ ..................... 5
2.3.1. I llumina Libr ary Preparation and Sequencing................................ .............. 5
2.3.2. I llumina Bioinformatics Workflow ................................ .............................. 5
2.4. L ineage Classification ................................ ................................ ............................... 6
2.5. Rules for L ineage Assignment ................................ ................................ .................. 6
3. COMPARI SON OF ION TORRENT AND ILLUM INA PLATFORMS ............................. 6
4. WORKFL OW AND DATA ENTRY AND STORAGE................................ ....................... 9
4.1. Workflow and LIMS Upload Process ................................ ................................ .......9
4.2. Data Repository ................................ ................................ ................................ ......... 9
5. REFERENCES ................................ ................................ ................................ .................... 10
LIST OF TABLES
Table 1. SARS -CoV -2 Lineage Assignments of PCR Positive NP Swabs
from Ion Torrent and Illumina NGS Platforms ................................ .......... 7
LIST OF FIGURES
Figure 1. Workflow for Capturing SARS -CoV -2 Sample Sequence Data ................ 9
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
FDA-CBER-2021-5683-1072367
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 31.BACKGROUND
1.1. Targeted NGS for SARS- CoV -2
Severe acute respiratory syndrome coronavirus 2 (SARS -CoV -2) is the virus that causes
coronavirus disease 2019 (COVID -19), the illness responsible for the COVID -19 pandemic.
SARS -CoV -2 is a positive -sense single -stranded RNA virus approximately 30,000 bases
inlength .Given the opportunity , RNA viruses have a propensit y to evolve - hundre ds of
thousands of distinct SARS -CoV -2viral lineages have been identified using next generation
sequencing (NGS) platforms. S ome of these are now classified as either variants of concern
(VOC) or variants of interest (VOI) based on disease characteristic s, epidemiology , and other
factors .NGS provides an effective, unbiased way to identify new coronavirus variants . This
information is valuable to understand viral evolution, geographic distribution, and
transmission. The sequence data may be used to inf orm public health decisions, mitigate
viral spread, and/or formulate recommendations for next generation vaccine s.
The viral content in clinical specimens that are PCR positive in SARS -CoV -2 molecular
diagnostic assay s can be highl y variable. In addition, the composition of nucleic acid
extracted from PCR positive clinical specimens is complex, including contributions from
thehuman host and microbes in addition to SARS -CoV -2. For these reasons ,a shotgun
metagenomics approach for determinatio n of SARS -CoV -2 genome sequence from clinical
specimens is not effective . Amplicon -based approaches that enrich for SARS -CoV -2 content
prior to NGS library construction are a necessary alternative. Following random primed
cDNA s ynthesis of total RNA in t he nucleic acid extracted from a clinical specimen,
selective SARS -CoV -2 enrichment is achieved by multiplex PCR using oligonucleotide sto
generate amplicons that are tiled across the viral genome. A library of these PCR amplicons
is constructed for each sample and then sequenced using NGS platforms.
1.2.Objective
The objective of this document is to describe the process used to determine the viral
sequence and lineage in SARS -CoV -2 RT -PCR (Cepheid) positive specimens from clinical
study C4591001.1Determination of SARS -CoV -2 lineage isan exploratory study objective .
Capturing SARS -CoV -2 sequence information from a SARS -CoV -2 PCR positive specimen
can be achieved using different sequencing platforms, two of which are manufactured by Ion
Torrent and Illumina.
2.MATERIALS AND METHOD S
2.1.Nucleic A cid Extraction
Midturbinate swabs that are positive in the SARS -CoV -2 RT -PCR (Cepheid) assay at either
the N or E gene target are advanced for viral sequence. Nucleic acid extractio n is performed
using the MagMAX ™Viral/Pathogen Ultra Nucleic Acid Isolation Kit processed on a
KingFisher Flex or KingFisher Presto.
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
FDA-CBER-2021-5683-1072368
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 42.2. Ion Torrent
2.2.1. Ion Torrent L ibrary P reparation and Sequencing
The SARS- CoV -2 viral genome sequencing is performed manuall y using the Ion
AmpliSeq™ technology and the Ion GeneStudio ™S5 plus sy stem.
The Ion AmpliSeq ™SARS -CoV -2 Research Panel consists of two primer pools that
target 237PCR amplicons specific to SARS -CoV -2 and 5 human expression controls.
Oligonucleotide primers based upon available SARS -CoV -2 nucleotide sequences direct
theamplification of the viral genome with amplicon lengths of 125 -275 bp. Thepanel
provides greater than 99% coverage of the SARS -CoV -2 genome (~30 kb), achieving
detect ion limits as low as 20 viral copies.
Briefl y, SARS -CoV -2 viral RNA content in the nucleic acid purified from swabs i s
quantified using TaqMan ™ 2019 -nCoV Assay Kit v1, the TaqMan™ 2019- nCoV
Control Kit v1, and the TaqPath ™ 1-Step RT -qPCR Master Mix, CG to determine the
optimal number of target amplification cy cles. cDNA is synthesized with the SuperScript
VILO cDNA s ynthesis Kit. Libraries are prepared using the Ion AmpliSeq™ library kit plus2
and Ion AmpliSeq™ SARS -CoV -2 research assay panel according to the manufacturer’s
instructions.3The prepared library under goes template preparation with the Ion C hef
according to the manufacturer’s instructions.4The enriched templates are then loaded onto
an Ion 530 chip for semiconductor sequencing on the I on Ge neStudio ™S5 plus se quencer
according to the manufacturer’s instructions.4
2.2.2. Ion Torrent B ioinformatics Workflow
Raw sequencing reads generated b y the Ion Torrent sequencer are quality and adaptor
trimmed by Ion Torrent Suite and the resulting reads are then mapped to the complete
genome of the SARS -CoV -2 Wuhan- Hu-1 isolate (GenBank accession number
MN908947.3) using TMAP 5.14.0.
Variant calling is carried out with the Torrent Variant Caller using the BAM file from the
mapping of the cleaned sequence reads to the reference sequence of SARS -CoV -2. Samples
not achieving mean sequence coverage of uniformity do not receive a
lineage assignment and ar e designated as indeterminant (IND). Samples designated as IND
are submitted for NGS using the Illumina platform ( Section 2.3).
Allidentified SARS -CoV -2 nucleotide variants areannotated using the software.5
For each of the variants identified, the output consists of: 1) the nucleotide of the reference at
each position and the alternative sequence, 2) the codon of the ref erence and the alternative
codon if the nucleotide variant results in a nons ynonymous substitution , and 3) the nucleotide
translation and information about the mutation (synony mous, missense plus deletions).
For each sample, a whole -genome sequence i s gene rated in a FASTA file using IRMA
(Iterative Refinement Meta -Assemble) and iterative optimiz ation of read gathering and
assembly .
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
(b) (4)
(b) (4)
(b) (4)
(b) (4)
FDA-CBER-2021-5683-1072369
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 52.3.Illumina
2.3.1. Illumina L ibrary P reparation and Sequencing
Briefl y, nucleic acid purified from swabs is digested with DNase using the Invitrogen
TURBO DNA -free™ Kit (AM1907) followed b y RNA purification using the Qiagen
RNeas y MinElute Cleanup Kit (74204). Initially performed manuall y, these steps (DNase
digestion and subsequent RNA purification) have since been programmed as additiona l steps
at the end of the primary RNA extraction using the KingFisher Presto ( Section 2.1).
Synthesis of cDNA is performed using ran dom sequence primers according to the AmpliSeq
for Illumina On -Demand, Custom and Community Panels Reference Guide.6
The cDNA is used as template to specificall y enrich for SARS -CoV -2 content by PCR. The
AmpliSeq for I llumina SARS -CoV -2 panel of PCR prim ers is a 2- pool design, containing a
total of 247 amplicons/primer pairs (Pool 1: 125 amplicons, Pool 2: 122 amplicons). These
include 242 SARS- CoV -2 viral -specific targets and 5 human gene expression controls . PCR
products range in size from 125- 275 bps in length and cover 99% of the viral genome and
all potential lineages of the virus. Oligonucleotide primers directing the amplification of the
viral amplicons are based upon available SARS -CoV -2 nucleotide sequences. Universal
Next G eneration Sequencing Adaptors are ligated to the ends of the SARS- CoV -2 AmpliSeq
amplicons. The amplicon libraries are purified with magnetic beads and loaded to a flow cell
for sequence determination using the Illumina NextSeq instrument according to the
manufacturer’s instructions.
2.3.2. Illumina B ioinformatics Workflow
A FASTQ file for each sample containing sequence data from the clusters that pass filter is
imported into CL C Genomic Workbench version 20.0.3. Any read with header information
that has a “fai led” flag due to a poor- quality score is removed during the file import process
by CLC Genomics Workbench. Primary reads are then aligned to the SARS -CoV -2
Wuhan -Hu-1 reference genome (Genbank accession MN908947.3) using the “Map Reads to
Reference” funct ion. Consensus SARS -CoV -2 sequence is generated b y using the “Extract
Consensus Sequence” function The
coverage of the Spike gene is verified b y using the “QC for Targeted Sequencing” function.
Only isolates that have coverage at each nucleotide position across the entire
Spike gene are advanced for lineage assignment. Single nucleotide variants (SNV) are called
using the “Low Frequency Variant Detection” function
Samples that do not meet these acceptance criteria are
considered indeterminant (IND) and not advanced for lineage assignment. Samples
designated as IND are submitted for NGS using the alternative Ion Torrent platform
(Section 2.2).
Nucleotide sequence coding for the Spike protein isextracted and translated into amino acid
sequence followed b y align ment to the Spike amino acid sequence of the Wuhan -Hu-1 strain
(Genbank accession number MN908947.3) using the CLC Genomics Workbench.
Non-synony mous substitutions relative to the Wuhan -Hu-1 Spike sequence are recorded.
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
(b) (4)
(b) (4)
(b) (4)
(b) (4)
FDA-CBER-2021-5683-1072370
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 62.4.Lineage Classification
SARS -CoV -2 lineage assignment for data from both I on Torrent and Illumina NGS
platforms is based on P angolin software,7which runs a multinomial logistic regression
model trained against lineage a ssignments based on isolate data from GI SAID,a global
science initiative established in 2008 that provides open- access to genomics data of influenza
virus and SARS- CoV -2.8
Viral lineage isdetermi nedfor samples sequenced using either the Ion Torrent or I llumina
NGS platform s, and in some cases using both platforms .
2.5.Rules for Lineage Assignment
SARS -CoV -2 lineage designations aredetermined from NGS data using either the Ion
Torrent or Illumina sequencing platforms. If acceptance criteria described in Section s 2.2.2
and Section 2.3.2 are met, a SARS- CoV -2 lineage is assigned and entered into the database.
If acceptance criteria are not met using the initial sequencing platform the result is considered
indeterminant (IND), and the sample is routed for assay using the alternative platform. If the
NGS determined using the alternative platform meets acceptance criteria, a SARS -CoV -2
lineage is assigned and entered into the database. If the NGS result is considered IND using
both platforms (ie, data do not meet acceptance criteria for either platform) and the Cepheid
RT-PCR Ct value for that sample is ≤ 34 for either the N or E gene target , the sample is
entered into the database as IND.
A sample is designated QNS (quantity not sufficient) if,
1.NGS quality does not meet the respective acceptance criteria (ie, sample is considered
IND) for both the Ion Torrent ( Section 2.2.2 ) and Illumina ( Section 2.3.2 ) platforms, and
2. T he sample has a Cepheid RT -PCR Ct value gr eater than 34 for both the N andE gene
target.
3. COMPARISON OF ION TORRENT AND ILLUMINA P LATFORMS
An initial group of Cepheid RT- PCR positive swab specimens from clinical study
C4591001 were processed as described above, using each of the two NGS platforms to
evaluate their suitability in determination of SARS -CoV -2 NGS and viral lineage. Selected
RT-PCR negative swabs were also included as negative controls. Sample barcodes,
Ctvalues for the N and E gene amplicons from the Cepheid PCR assay, an d SARS -CoV -2
lineage assignments determined from nucleotide sequence collected using the Ion Torrent
and I llumina NGS platforms are listed i n Table 1. The Illumina NGS did not meet
acceptance criteria for of the samples (22.5%). Several of these had very high Cepheid
PCR Ct values (ie, low SARS- CoV -
2 RNA content). Lineage assignments could not be
made from Ion Torrent NGS for of the samples (14.1%); of these were also
designated as IND from Illumina NGS.
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
(b) (4)
(b) (4)
(b) (4)
(b) (4)
(b) (4)
(b) (4)
FDA-CBER-2021-5683-1072371
FDA-CBER-2021-5683-1072372
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 8Table 1.SARS -CoV -2 Lineage Assignments of PCR Positive NP Swabs from Ion
Torrent and Illumina NGS Platforms
Sample Barcodes Cepheid E (Ct)aCepheid N (Ct) *Ion Torrent Lineage Illumina Lineage
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
(b) (4)
FDA-CBER-2021-5683-1072373
FDA-CBER-2021-5683-1072374
PF-07302048 : SARS -CoV- 2 Whole Genome Sequencing Data Collection and A nalysis Guidelines
VR-VTN -10436 , Ver. 1.0
PFIZER CONFIDENTIAL
Page 105.REFERENCES
1VR-VR-10080. Report on Method Validation of a Cepheid Xpert® Xpress PCR Assay
to Detect SARS -CoV -2.
2ThermoFisher. Ion AmpliSeq™ Library Kit Plus USER GUIDE. Publication
MAN0017003 version C.0.
3Ion AmpliSeq™ SARS -CoV2 Research Panel, Instructions – for use on an I on
GeneStudio™ S5 Series Sy stem. Publication MAN0019277, version B.0 .
4ThermoFisher. Ion 510 ™& Ion 520 ™& Ion 530™ Kit –Chef USER GUIDE.
Publication MAN0016854, version F.0.
5
6Illumina. AmpliSeq for Illumina On -Demand, Custom and Community Panels
Reference Guide Document 1000000036408 , v09 May 2020. Accessed from:
https://support.illumina.com/content/dam/illumina -
support/documents/documentation/chemistry _documentation/ampliseq -for-
illumina/ampliseq -for-illumina -custom -and-community -panels -reference -guide -
1000000036408 -09.pdf
7Pangolin, https://github.com/cov -lineages/p angolin
8https://www.gisaid.org/
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
(b) (4)
FDA-CBER-2021-5683-1072375
Document Approval Record
Document Name:
Document Title: !"!"!#!$%
%#%
Signed By: Date(GMT) Signing Capacity
&!'(
)
*+*, "-..-/!
01-!"-'2!
)
+**3 4!..-/!
0'5
)
+** "-..-/!
6!%'7!
)
*
*
4!..-/!
)
*+*3
7!!'!-- )
+**
8!!-..-/!
090177e19732889a\Approved\Approved On: 04-Jun-2021 15:42 (GMT)
(b) (6)
(b) (6)
FDA-CBER-2021-5683-1072376