Translational Medicine

EFTA00315100 Dataset 9 14 pages Download original PDF Download as text
Ience Translational Medicine AAAS Rapid Whole-Genome Sequencing for Genetic Disease Diagnosis in Neonatal Intensive Care Units Carol Jean Saunders et at Sci Trans! Med 4, 154ra135 (2012); DOI: 10.1126/scitranslmed.3004041 Editor's Summary Speed Heals The waiting might not be the hardest part for families receiving a diagnosis in neonatal intensive care units (NICUs), but it can be destructive nonetheless. While they wait on pins and needles for their newborn baby's diagnosis, parents anguish, nurture false hope, wrestle with feelings of guilt—and all the while, treatment and counseling are delayed. Now, Saunders et at describe a method that uses whole-genome sequencing (WGS) to achieve a differential diagnosis of genetic disorders in 50 hours rather than the 4 to 6 weeks. Many of the -3,500 genetic diseases of known cause manifest symptoms during the first 28 days of life, but full clinical symptoms might not be evident in newborns. Genetic screens performed on newborns are rapid, but are designed to unearth only a few genetic disorders, and serial gene sequencing is too slow to be clinically useful. Together, these complicating factors lead to the administration of treatments based on nonspecific or obscure symptoms, which can be unhelpful or dangerous. Often, either death or release from the hospital occurs before the diagnosis is made. 0 The new WGS protocol cuts analysis time by using automated bioinformatic analysis. Using their newly 2 developed protocol, the authors performed retrospective 50-hour WGS to confirm, in two children, known molecular diagnoses that had been made using other methods. Next, prospective WGS revealed a molecular diagnosis of a u- BRAT1-related syndrome in one newborn; identified the causative mutation in a baby with epidermolysis bullosa: ruled out the presence of defects in candidate genes in a third infants; and, in a pedigree, pinpointed BCL9L as a new co recessive gene (HTX6) that gives rise to visceral heterotaxy —the abnormal arrangement of organs in the chest and $6 abdominal cavities. WGS of parents or affected siblings helped to speed up the identification of disease genes in the cif prospective cases. These findings strengthen the notion that WGS can shorten the differential diagnosis process and E quicken to move toward targeted treatment and genetic and prognostic counseling. The authors note that the speed and cost of WGS continues to rise and fall, respectively. However, fast WGS is clinically useful when coupled with fast c and affordable methods of analysis. -a E E -§ a A complete electronic version of this article and other services, including high-resolution figures, can be found at: http://stm.sciencemag.org/content/4/154/154ra135.full.html Supplementary Material can be found in the online version of this article at: http://stm.sciencemag.org/content/suppl/2012/10/01/4.154.154ra135.DC1.html Related Resources for this article can be found online at: http://stm.sciencemag.org/content/scitransmed/5/194/194cm5.full.html http://stm.sciencemag.org/content/scitransmed/5/194/194ed10.full.html http://stm.sciencemag.org/content/scitransmed/5/198/198ed12.full.html http://www.sciencemag.org/content/sci/342/6155/197.full.html Information about obtaining reprints of this article or about obtaining permission to reproduce this article in whole or in part can be found at: http://www.sciencemag.org/about/permissions.dtl Science Translational Medicine (print ISSN 1946-6234; online ISSN 1946-6242) is published weekly, except the last week in December. by the American Association for the Advancement of Science, 1200 New York Avenue NW, Washington, DC 20005. Copyright 2012 by the American Association for the Advancement of Science; all rights reserved. The title Science Translational Medicine is a registered trademark of AAAS. EFTA00315100 RESEARCH ARTICLE DIAGNOSTICS Rapid Whole-Genome Sequencing for Genetic Disease Diagnosis in Neonatal Intensive Care Units Carol Jean Saunders,1,2,3,4,5* Neil Andrew MiIler,l'2'4* Sarah Elizabeth Soden," 4* Darrell Lee Dinwiddie," 31"* Aaron Noll,' Noor Abu Alnadi,4 Nevene Andraws, 3 Melanie LeAnn Patterson," Lisa Ann Krivohlavek," Joel Ferns,' Sean Humphray,' Peter Saffrey,' Zoya Kingsbury,' Jacqueline Claire Weir,' Jason Betley," Russell James Grocock,' Elliott Harrison Margulies,' Emily Gwendolyn Farrow,' Michael Artman,2'4 Nicole Pauline Safina," Joshua Erin Petrikin," Kevin Peter Hall,' Stephen Francis Kingsmore'•2,3,4,st Monogenic diseases are frequent causes of neonatal morbidity and mortality, and disease presentations are often undifferentiated at birth. More than 3500 monogenic diseases have been characterized, but clinical testing is avail- able for only some of them and many feature clinical and genetic heterogeneity. Hence, an immense unmet need exists for improved molecular diagnosis in infants. Because disease progression is extremely rapid, albeit hetero- geneous, in newborns, molecular diagnoses must occur quickly to be relevant for clinical decision-making. We de- scribe 50-hour differential diagnosis of genetic disorders by whole-genome sequencing (WGS) that features automated bioinformatic analysis and is intended to be a prototype for use in neonatal intensive care units. Ret- rospective 50-hour WGS identified known molecular diagnoses in two children. Prospective WGS disclosed potential molecular diagnosis of a severe G22-related skin disease in one neonate; BRAT/-related lethal neonatal rigidity and multifocal seizure syndrome in another infant identified Ba9L as a novel, recessive visceral heterotaxy gene (H7X6) in a pedigree; and ruled out known candidate genes in one infant. Sequencing of parents or affected siblings expedited the identification of disease genes in prospective cases. Thus, rapid WGS can potentially broaden and foreshorten differ- ential diagnosis, resulting in fewer empirical treatments and faster progression to genetic and prognostic counseling. INTRODUCTION Genomic medicine is a new, structured approach to disease diagnosis and management that prominently features genome sequence infor- mation (1). Whole-genome sequencing (WGS) by next-generation sequencing (NGS) technologies has the potential for simultaneous, comprehensive, differential diagnostic testing of likely monogenic ill- nesses, which accelerates molecular diagnoses and minimizes the du- ration of empirical treatment and time to genetic counseling (2-7). Indeed, in some cases, WGS or exome sequencing provides molecular diagnoses that could not have been ascertained by conventional single- gene sequencing approaches because of pleiotropic clinical presenta- tion or the lack of an appropriate molecular test (7-9). Neonatal intensive care units (NICUs) are especially suitable for early adoption of diagnostic WGS because many of the 3528 mono- genic diseases of known cause are present during the first 28 days of life (10). In the United States, more than 20% of infant deaths are caused by congenital malformations, deformations, and chromosomal abnormalities that cause genetic diseases (11-13). Although this pro- portion has remained unchanged for the past 20 years, the precise prevalence of monogenic diseases in NICUs is poorly understood be- cause ascertainment rates are low. Serial gene sequencing is too slow to be clinically useful for NICU diagnosis. Newborn screens, while 'Center for Pediatrocancenc Medidne Children's Mercy Hospital• Kansas City,M0 69103, USA /Department of Pediatric,. Children's Mercy Hospital. Kansas City. MO 6910. IA& 'Department of Pathology. Children's Mercy Hosptat Kansas City. M069103. USA °School of Medicine. University of Missoun-Kansas Cny. Kansas City. MO 6410& USA. sUni.ersity of Kansas Medical Center. KansisCny. KS66160, USA'1Ilumina Inc,Chesterfad Research Park Little Chesterford. CBIO I XL Essex UK 'These authors contributed equally to this work tic, whom correspondence should be addressed. E-mail stkingsrnore@cmhedu 0 rapid, identify only a few genetic disorders for which inexpensive tests and cost-effective treatments exist (14, 15). Further complicating diag- nosis is the fact that the full clinical phenotype may not be manifest • in newborn infants (neonates), and genetic heterogeneity can be im- mense. Thus, acutely ill neonates with genetic diseases are often dis- charged or deceased before a diagnosis is made. As a result, NICU treatment of genetic diseases is usually empirical, may lack efficacy, may be inappropriate, or may cause adverse effects. NICUs are also suitable for early adoption of genomic medicine because extraordinary interventional efforts are customary and inno- vation is encouraged. Indeed, NICU treatment is among the most cost-effective of high-cost health care, and the long-term outcomes of most NICU subpopulations are excellent (16-18). In genetic diseases for which treatments exist, rapid diagnosis is critical for timely delivery of interventions that lessen morbidity and mortality (14-17, 19, 20). For neonatal genetic diseases without effective therapeutic interven- tions, of which there are many (21), timely diagnosis avoids futile inten- sive care and is critical for research to develop management guidelines that optimize outcomes (22). In addition to influencing treatment, neo- natal diagnosis of genetic disorders and genetic counseling can spare parents diagnostic odysseys that instill inappropriate hope or perpetuate needless guilt Two recent studies exemplify the diagnostic and therapeutic uses of NGS in the context of childhood genetic diseases. WGS of fraternal twins concordant for 3,4-dihydroxyphenylalanine (dopa)-responsive dystonia revealed known mutations in the sepiapterin reductase (SPR) gene (3). In contrast to other forms of dystonia, treatment with 5-hydroxytryptamine and serotonin reuptake inhibitors is beneficial in patients with SPR defects. Application of this therapy in appropriate cases resulted in clinical improvement. Likewise, extensive testing a) C 0 rn til E 8 C tb O N E E 8 wiwSdenceTranslatIonalMedlamorg 3 October 2012 Vol 4 Issue 154 154ral 3S 1 EFTA00315101 RESEARCH ARTICLE failed to provide a molecular diagnosis for a child with fulminant pan- colitis (extensive inflammation of the colon) (8), in whom standard treatments for presumed Crohn's disease—an inflammatory bowel disease—were ineffective. NGS of the patient's exome, together with confirmatory studies, revealed X-linked inhibitor of apoptosis (XIAP) deficiency. The treating physicians had not entertained this diagnosis because X1AP mutations had not previously been associated with co- litis Hemopoietic progenitor cell transplant was performed, as indi- cated for XIAP deficiency, with complete resolution of colitis. last, for —3700 genetic illnesses for which a molecular basis has not yet been established (10), WGS can suggest candidate genes for functional and inheritance -based confirmatory research (23). The current cost of research-grade WGS is $7666 (24)-which is similar to the current cost of commercial diagnostic dideoxy sequencing of two or three disease genes. Within the context of the average cost per day and per stay in a NICU in the United States (13), WGS in care- fully selected cases is acceptable and even potentially cost-saving (3-7). However, the turnaround time for interpreted WGS results, such as that of dideoxy sequencing, is too slow to be of practical use for NICU diagnoses or clinical guidance (typically -4 to 6 weeks) (2-4). Here, we report a system that permits WGS and bioinfonnatic analysis (largely automated) of suspected genetic disorders within 50 hours, a time frame that appears to be promising for emergency use in level 3 NICUs. RESULTS Symptom- and sign-assisted genome analysis (SSAGA) is a new din- icopathological correlation tool that maps the clinical features of 591 well-established, recessive genetic diseases with pediatric presentations (table SI) to corresponding phenotypes and genes known to cause the symptoms (2, 10). SSAGA was developed for comprehensive auto- mated performance of the following two tasks: (i) WGS analyses re- stricted to a superset of gene-associated regions relevant to clinical presentations, in accord with published guidelines for genetic testing in children (25-28), and (ii) prioritization of clinical information to assist in the interpretation of WGS results. SSAGA has a menu of 227 clinical terms arranged in nine symptom categories (fig. SI). Stan- dardized clinical terms (29) have been mapped to 591 genetic diseases on the basis of authoritative databases (10, 30) and expert physician reviews. Each disease gene is represented by an average of 8 terms and at most 11 terms (minimum, I term, 15 disease genes; maximum, 11 terms, 3 disease genes). To validate the feasibility of automated matching of clinical terms to diseases and genes, we entered retrospectively the presenting fea- tures of 533 children who have received a molecular diagnosis at our institution [Children's Mercy Hospital (CMH), Kansas City, MO] within the last 10 years into SSAGA. Sensitivity was 99.3% (529), as determined by correct disease and affected gene nominations. Failures induded a patient with glucose-6-phosphate dehydrogenase deficiency who presented with muscle weakness [which is not a feature men- tioned in authoritative databases (10, 30)]; a patient with Janus kinase 3 mutations who had the term "respiratory infection" in his medical records rather than "increased susceptibility of infections," which is the description in authoritative databases; and a patient with cystic fibrosis who had the term "recurrent infections" in his medical records rather than "respiratory infections," which is the description in au- thoritative databases. SSAGA nominated an average of 194 genes per patient (maximum, 430; minimum, 5). Thus, SSAGA displayed sufficient sensitivity for the initial selection of known, recessive candi- date genes in children with specific clinical presentations. Rapid WGS To assess our ability to recapitulate known results, we performed rapid WGS retrospectively on DNA samples from two infants with molec- ular diagnoses that had previously been identified by clinical testing. Then, to assess the potential diagnostic use of rapid WGS, we prospec- tively performed WGS in five undiagnosed newborns with clinical presentations that strongly suggested a genetic disorder as well as their siblings. Automation of the five main components of WGS as well as bioinformatics-based gene-variant characterization and clinical inter- pretation, all in an integrated workflow, made possible —50-hour time to differential molecular diagnosis of genetic disorders (Fig. 1). Specifically, sample preparation for WGS was shortened from 16 to 4.5 hours, while a physician simultaneously entered into SSAGA clin- ical terms that described the neonates' illnesses (fig. SI). For each sample, rapid WGS [2 x 100 base pair (bp) reads, including on-board cluster generation and paired-end sequencing] was performed in a single run on the alumina HiSeq 2500 and took —26 hours. Base calling, genomic sequence alignment, and gene variant calling took —15 hours. The HiSeq 2500 runs yielded 121 to 139 gigabases (GB) of aligned sequences (34- to 4I-fold aligned genome coverage; 'fable I). Eighty-eight to 91% of 5 bases had >99.9% likelihood of being correct (quality score >30, using Illumina software equivalent to Phred) (31, 32). We detected 4.00 ± 0.20 million nudeotides that differed from the reference genome se- quence (variants) (mean ± SD) in nine samples, one from each of nine infants (Table 1). Analytical metrics In three samples, genome variants identified by 50-hour WGS were compared with those identified by deep targeted sequencing of either exons and 20 intron-exon boundary nucleotides of a panel of 525 re- cessive disease genes [Children's Mercy Hospital Diagnostic panel 1 (CMH-Dxl)] or the exome (Table 2). CMH-Dx1 comprised 8813 exonic and intronic targets, totaling 2.1 million nucleotides (table SI) (2, 33). The exome and CMH-Dx1 methods, which used Illumina TruSeq enrichment and HiSeq 2000 sequencing, took —19 days. In contrast, rapid WGS did not use target enrichment, was performed with the HiSeq 2500 instrument, and took —50 hours. Samples CMH064, UDT002, and UDTI73 were sequenced using these three methods, and variants were detected with a single alignment method [the Genomic Short-read Nucleotide Alignment Program (GSNAP)] (34) and variant caller [the Genome Analysis Tool Kit (GATK)] (35). Rapid WGS detected —96% of the variants identified by a target en- richment method and —99.5% of the variants identified by both methods had identical genotypes (Table 2), indicating that rapid WGS is highly concordant with established clinical sequencing methods (33). In contrast, analysis of the rapid WGS data set from sample CMH064 with three different alignment and variant detection methods [GSNAP/GATK, the alumina CASAVA alignment tool, and the Burrows-Wheeler Alignment (BWA) tool] revealed surprising dif- ferences between the variants detected. Only —80% of the variants de- tected using GATK/GSNAP or BWA were also detected with CASAVA (Table 2 and table S2) (36-41). This suggests that additional studies will be needed to define optimal alignment methods for dinical sequencing. asas 2 .o ts. ,c2) E of C N E E -0 cu 8 wwwSdenceTranslatIonalMediclnenrg 3 October 2012 Vol 4 Issue 154 154ra135 2 EFTA00315102 RESEARCH ARTICLE Obtain consent and blood sample Prepare sequencing library Enter clinical findings into SSAGA S HiSeq 2500 2 x 100 bp sequencing CASAVA base calling RUNES variant annotation SSAGA-delimited variant analysis and interpretation Verbal interim report of diagnosis pending CLIA confirmation Fig. 1. STAT-Sea. Summary of the steps and timing of STAT-Seq, result- ing in an interval of 50 hours between consent and delivery of a pre- liminary, verbal diagnosis. t, hours. Nevertheless, there was good concordance between the genotypes of variants detected by rapid WGS (using the HiSeq 2500 and CASAVA) and targeted sequencing (using exome enrichment, the HiSeq 2000, and GATKIGSNAP)-99.5% (UDT002), 99.9% (UDT173), and 99.7% (CMH064) (Table 2)—further indicating that rapid WGS is highly con- cordant with an established genotyping method (33). In subsequent studies, the rapid WGS technique used CASAVA for alignment and variant detection. Genomic variants were characterized with respect to functional consequence and zygosity with a new software pipeline [Rapid Understanding of Nucleotide variant Effect Software (RUNES), fig. S21 that analyzed each sample in 2.5 hours. Samples contained a mean of 4.0D ± 0.20 million (SD) genomic variants, of which a mean of 1.87 ± 0M9 million (SD) were se-striated with protein-encoding genes (Table 1). Less than I% of these variants (mean, 10,848 ± 523 SD) were also of a functional class that could potentially be disease causative (Fable 1) (25-27). Of these, —14% (mean, 1530 ± 518 SD) had an allele fiequen- cy that was sufficiently low to be a candidate for being causative in an uncommon disease (<1% allele frequency in 836 individuals sequenced d at CMH) (42). Last, of these, —71% (mean, 1083 ± 240 SD) were also of a functional class that was likely to be disease causative [American i• College of Medical Genetics (ACMG) categories Ito 31 (Table 1). This 2 set of variants was evaluated for disease causality in each patient, with priority given to variants within the candidate genes that had been Li- nominated by an individual patient presentation. Es Retrospective analyses ° Patient UDT002 was a male who presented at 13 months of age with na hypotonia, developmental regression. Brain magnetic resonance imag- 5 ing (MRI) showed diffuse white matter changes suggesting leukodys- trophy. Three hundred fifty-two disease genes were nominated by .20 one of the three clinical terms hypotonia, developmental regression, or 01 leukodystrophy, 150 Aise-ase genes were nominated by two terms; and 9 disease genes were nominated by all three terms (table S3). Only E 16 known pathogenic variants had allele frequencies in dbSNP and the 2 CMH cumulative database that were consistent with uncommon dis- ease mutations. Of these, only two variants mapped to the nine can- didate genes; the variants were both compound heterozygous (verified cog, by parental testing) substitution mutations in the gene that encodes the t a subunit of the lysosomal enzyme hexosaminidase A [HEXA Chr 15:72,641,417c>C (gene symbol, chromosome number, chromosome coordinate, reference nucleotide > variant nucleotide), c.986+3A>G (transcript coordinate, reference nucleotide, variant nucleotide), and Chr15:72,640,388C>T, c.1073+1G>A1. The c.986+3A>G alters a 5' exon-flanking nucleotide and is a known mutation that causes Tay-Sachs disease ('15D), a debilitating lysosomal storage disorder [Online Mendelian Inheritance in Man (OMIM) number 2728001. The variant had not previously been observed in our database of 651 individuals or dbSNP, which is relevant because mutation databases are contaminated with some common polymorphisms, and these can be distinguished from true mutations on the basis of allele frequency (33). The c1073+1G>A variant is a known l'SD mutation that affects an exonic splice donor site (dbSNP r576173977). The variant has been observed only once before in our database of 414 samples, which is consistent with an allele frequen- cy of a causative mutation in an orphan genetic disease. Thus, the known diagnosis of 'Est) was confirmed in patient UDT002 by rapid WGS. Patient UDT173 was a male who presented at 5 months of age with developmental regression, hypotonia, and seizures. Brain MRI showed www.5cienceTranslationalMedlciae.org 3 October 2012 Vol 4 Issue 154 154ral 35 3 EFTA00315103 RESEARCH ARTICLE dysmyelination, hair shaft analysis revealed pill ford (kinky hair), and serum copper and ceruloplasmin were low. On the basis of this clinical presentation, 276 disease genes matched one of these clinical terms and 3 matched three terms (table S4). 'There were no previously reported disease-causing variants in these 276 genes. However, five of the candi- date genes contained either variants of a type that is expected to be disease-causing based on their predicted functional consequence or missense variants of unknown significance (VUS). One of these var- iants was in a gene that matched all three clinical terms and was a hemizygous substitution mutation in the gene that encodes the a poly- peptide of copper-transporting adenosine triphosphatase (ATP7A Chr )C77,271,307C>T, c.2555C>T, i, aberrant forms of which are known to cause Menkes disease, a copper-transport disorder. This variant—new to our database and dbSNP—specified a nonconserva- tive substitution in an amino acid that was highly conserved across species and had deleterious SIFT (Sons Intolerant From Tolerant sub- stitutions), PolyPhen2 (Polymorphism Phenotyping), and BLOSUM (B1nrlec SUbstitution Matrix) scores. The known diagnosis of Menkes disease (OMIM number 309400) was recapitulated. As a further assess- ment of the reliability of variant detection of rapid WGS, samples UDT002 and UDTI73 were aligned to the reference genome with three different alignment methods. The causative variants were recovered with each method. Prospective analyses Mutations in 35 genes can cause generalized, erosive dermatitis of the type found in CMH064 (table S5). The severe phenotype, negative family history, and absence of consanguinity suggested dominant de novo or recessive inheritance. No known pathogenic mutations were identified in the candidate genes that had low allele frequencies in the CMH cumulative genome and exome sequence database and similar public datakcec Average coverage of the genomic regions corresponding to the candidate genes was 38.9-fold, and 98.4% of candidate gene nucleotides had >I6x high-quality coverage (sufficient to rule out a n Table 1. Sequencing, alignment, and variant statistics of nine samples analyzed by rapid WGS. ACMG category 1 to 4 variants are a subset of gene- (5 associated variants. L'" ets n 2 or High- ACMG ACMG ACMG Candidate Run Sampletimenor Mitochondria )genome Nuclear Sequence quality genomeGene-categories categories categories 1 to 4 1 to 3 Candidate gene Candidate c ° (hours) variants variants variants (%) variants frequency frequency 1 variants <1% <1%0) to UDT002 255 133 91 33 4,014,761 1,888,650 10,733 1,989 1,330 352 (9) 2 0 E UDT173 255 139 89 40 3,977,062 1,859,095 10,501 2,190 1,296 347 (3) 0 I C CD CMH064 26.6 121 88 41 3,985,929 1,869,515 10,701 1.884 1,348 35 0 2 to CMH076 25.7 134 88 34 4,498,146 2,098,886 11,891 2,552 1,351 89 0 CMH172 265 113 91 39 3,759,165 1,749,868 10,135 1,456 982 174 0 N CMH184 265 137 90 37 3,921,135 1,840,738 10,883 1,168 833 12 0 0 E 2 CMH185 40 117 93 37 3,922,736 1,831,997 10,810 1,164 840 14 0 0 CMH186 255 113 93 37 3,933,062 1,827,499 10,713 1,202 868 14 § CMH2O2 40 116 93 39 3,947,053 1,849,647 10,805 1,283 901 C C 0 0 Table 2. Variants and genotypes. Comparisons of variants and genotypes obtained In three samples using three target enrichment methods, two se- quencing methods, and two alignment methods. The SO hour WGS (STAT-Seci) was not enriched and used HiSeq 2500 sequencing. CMHDx1 was erviched for 523 genes and Merl 2000 sequencing. Average coverage of target nucleo- tides indicates the average aligned sequence depth over the corresponding target panel. For WGS, the target is the genome; for come sequencing, the target is the exome and for CMH-Corl, the targets are 523 genes. Sample Target enrichment Sequencing method Alignment method Sequence (GB) Average coverage of target nucleotides Variants detected by rapid WGS Genotypes Identical to both methods (%) CMH064 Exorne HiSeq 2000 GATK/GSNAP 9.8 79 46,756 (96.0%) 99.4 None (WGS) HiSeq 2500 12.1 40 UDT173 CMH-Dxl HiSeq 2000 GATK/GSNAP 4.1 784 1539 (96.7%) 99.60 None (WGS) HiSeq 2500 13.9 46 UDT173 CMH-Dxl HiSeq 2000 GATK/GSNAP 4.1 784 1457 (83.0%) 99.9 None (WGS) HiSeq 2500 CASAVA 13.9 46 UDT002 CMH-Dxl HiSeq 2000 GATK/GSNAP 4.2 770 1341 (76.6%) 99.5 None (WGS) HiSeq 2500 CASAVA 13.3 44 www.ScienceTranslationalMedicine.org 3 October 2012 Vol 4 Issue 154 154ra135 4 EFTA00315104 RESEARCH ARTICLE heterozygous variant; table S6). Five candidate genes had 100% nu- cleotides with >16-fold high-quality coverage and, thus, lacked a known pathogenic mutation in an exon or within 20 nucleotides of the intron- exon boundaries. Eighteen of the candidate genes had >99% nucleotides with >16-fold high-quality coverage, and 31 had >95% nucleotides with at least this level of coverage. Furthermore, while 26 of the candidate genes had pseudogenes, paralogs, and/or repeat segments (table S6) that could potentially result in misalignment and variant miscalls, only 0.03% of target nucleotides had poor alignment quality scores. Among the 35 candidate genes nominated by the phenotype, two rare heterozygous VUS were detected in CMH064; however, dideoxy sequencing of both healthy parents exduded one, in the keratin 14 gene, as a de novo mutation. The exomes of both parents were sub- sequently sequenced, and variants were examined in the trio at length. Three likely de novo mutations with excellent sequence coverage were identified in disease-causing genes. Of these, one was a candidate gene for CMH064. It was an in-frame deletion of three nucleotides in GIB2, (NM_004004), which encodes the connedn 26 protein. The variant, c85_87de1, removes a highly conserved amino acid within the first transmembrane helix (43). Dideoxy sequencing confirmed it to be a de novo mutation. Dominant, de novo 6,1B2 mutations have been associated with severe neonatal lethal disorders of the skin, such as keratitis-ichthyosis-deafness syndrome (KIDS), that involve the suprabasilar layers of the epidermis (OMIM number 148210) (44). The phenotype of CMH064 was atypical for KIDS, and functional studies are in progress to determine causality definitively. Diagnoses suggested by the presentation in CMH076 were mito- chondrial disorders, organic acidemia, or pyruvate carboxylase defi- ciency. Together, 75 nuclear genes and the mitochondrial genome cause these diseases (table S7). A negative family history suggested re- cessive inheritance that resulted from compound heterozygous or hemi- zygous variants or a heterozygous de novo dominant variant. Rapid WGS excluded known pathogenic mutations in the candidate genes. One novel heterozygous VUS was found. However, de novo occur- rence of this variant was ruled out by exome sequencing of his healthy parents. No homozygous or compound heterozygous VUS with suit- ably low allele frequencies were identified in the known disease genes. Potential novel candidates included 929 nuclear genes that encode mitochondrial proteins but have not yet been associated with a genetic disease (45). Only one of these had a homozygous or compound het- erozygous VUS with an allele frequency in dbSNP and the CMH database that was sufficiently low to be a candidate for causality in an uncommon inherited disease. Deep exome sequencing of both parents excluded this variant and did not disclose any further poten- tially causal variants. A total of 174 genes are known to cause epilepsy of the type found in CMH 172 (table S8). A positive family history of neonatal epilepsy and evidence of shared parental ancestry strongly suggested recessive inheri- tance. No known disease-causing variants or homozygous/compound heterozygous VUS with low allele frequencies were identified in these genes, which largely excluded them as causative in this patient A genome- wide search of homozygous, likely pathogenic VUS that were novel in the CMH database and dbSNP disclosed a frame-shifting insertion in the BRCA)-associated protein required for ATM activation-I (BRAT), Chr 7:2,583,573.2,583,574insATCITCTC,c453_454insATCITCTC, . A literature search yielded a very recent study of BRAT) mutations in two infants with lethal, multifocal seizures, hyper- tonia, microcephaly, apnea, and bradycardia (OMINI number 614498) (46). Dideoxy sequencing confirmed the variant to be homozygous in CMH172 and heterozygous in both parents. Rapid WGS was performed simultaneously on proband CMHI84 (male), affected sibling (brother) CMHI85, and their healthy parents, CMH186 and CMH2O2. Twelve genes have been associated with the clinical features of the brothers (heterotaxy and congenital heart dis- ease table S9). Co-occurrence in two siblings strongly suggested reces- sive inheritance. No known disease-causing variants or homozygous/ compound heterozygous VUS with low allele frequencies were identi- fied in these genes. A genome-wide search of novel, homozygous/ compound heterozygous, likely pathogenic VUS that were common to the affected brothers and heterozygous in their parents yielded two nonsynonymous variants in the B cell CLL/ghc -like gene (BCL9L, Chr 11:118,772,350G>A,c.2102G>A, and Chr 11:118,774,140G>A, c.554C>T, i. Evidence supporting the candidacy of BCL9L for heterotaxy and congenital heart disease is presented below. 0 DISCUSSION 0 a's Genomic medicine, empowered by WGS, has been heralded as trans- g formational for medical practice (2, 4, 5, 47). Over the last several y years, the cost of WGS has fallen markedly, potentially bringing it LI- within the realm of cost-effectiveness for high-intensity medical prac- tice, such as occurs in NICUs (3, 8, 23, 24). Furthermore, experience IF has been gained with clinical use of WGS that has instructed initial ; guidelines for its use in molecular diagnosis of genetic disorders (9). However, a major impediment to the implementation of practical ge- nomic medicine has been time to result a) This limitation has always been a problem for diagnosis of genetic .0 disease& Time to result and cost have greatly constrained the use of 1.'1 serial analysis of single-gene targets by dideoxy sequencing, Hitherto, clinical use of WGS by NGS has also taken at least a month: Sample preparation has taken at least a day; clustering 5 hours; 2 x 100 nu- O cleotide sequencing 11 days; alignment, variant calling, and genotyping 1 day, variant characterization a week and clinical interpretation at least a week Although exome sequencing lengthens sample preparation by 113 several days, it decreases computation time somewhat and is less costly. t For use in acute care, the turnaround time of molecular diagnosis, 3 including analysis, must match that of medical decision-making, which ranges from Ito 3 days for most acute medical care. Herein, we de- scribed proof of concept for 2-day genome analysis of acutely ill neo- nates with suspected genetic disorders. Automating medicine Rapid WGS was made possible by two innovations. First, a widely used WGS platform has been modified to generate up to 140 GB of sequence in less than 30 hours (HiSeq 2500): Sample preparation took 4.5 hours, and 2 x 100 bp genome sequencing took 25.5 hours (Fig. 1). The total "hands-on" time for technical staff was 5 hours. Modifica- tions included a new flowcell design and faster imaging and chemistry. Previously, NGS has either lacked sufficient sequence quantity, quality, or read lengths for clinical use of WGS or been too slow for use in acute patient care. Rapid WGS generated -40-fold aligned genome coverage. The sequence quality was very similar to that obtained with its predeces- sor (HiSeq 2000), as determined by quality scores and alignment rates (48). Genotypes of nucleotide variants were >99.5% concordant with wwwSdenceTranslatIonalMedicintorg 3 October 2012 Vol 4 Issue 154 154ral 35 5 EFTA00315105 RESEARCH ARTICLE those of very deeply sequenced, partial exomes (33). The accuracy of the latter has been extensively benchmariced and is >99.9% (33). Second, we automated much of the onerous characterization of ge- nome variation and facilitated interpretation by restricting and prior- itizing variants with respect to allele frequency (42), likelihood of a functional consequence (25), and relevance to the prompting illness. Thus, rapid WGS, as described herein, was designed for prompt dis- ease diagnosis rather than carrier testing or newborn screening. SSAGA mapped the clinical features in ill neonates and children to disease genes. Thereby, analysis was limited only to the parts of the genome relevant to an individual patient's presentation, in accord with guide- lines for genetic testing in children (25-28). This greatly decreased the number of variants to be interpreted. In particular, SSAGA caused most incidental (secondary) findings to be masked. In the setting of acute care in the NICU, secondary findings are anticipated to impede facile interpretation, reporting, and communication with physicians and patients greatly (9, 49, 50). SSAGA also assisted in test ordering, permitting a broad selection of genes to be nominated for testing based on entry of the patients' clinical features with easy-to-use pull-down menus. The version used herein contains —600 recessive and mitochon- drial diseases and has a diagnostic sensitivity of 993% for those dis- orders. SSAGA is likely to be particularly useful in disorders that feature clinical or genetic heterogeneity or early manifestation of partial phenotypes because it maps features to a superset of genetic disorders. SSAGA needs to be expanded to encompass dominant disorders and to the full complement of genetic diseases that meet ACMG guidelines for testing rare disorders (such as having been reported in at least two unrelated families) (26). Although neonatal disease presentations are often incomplete, only one feature is needed to match a disease gene to a presentation. In cases for which SSAGA-delimited genome analysis was negative, such as CMH064 and CMH076, a comprehensive second- ary analysis was performed with limitation of variants solely to those with acceptable allele frequencies (42) and likelihood of a functional consequence (25). Nevertheless, secondary analysis was relatively facile, yielding about 1000 variants per sample. RUNES performed many laborious steps involved in variant char- acterization, annotation, and conversion to HGVS (Human Genome Variation Society) nomenclature in -2 hours. RUNES unified these in an automated report that contained nearly all of the information de- sirable for variant interpretation, together with a cumulative variant allele frequency and a composite ACMG categorization of variant pathogenicity (fig. S2). ACMG categorization is a particularly useful standard for prioritization of the likelihood of variants being causal (26). In particular, more than 75% of coding variants were of ACMG category 4 (very unlikely to be pathogenic). Removal of such variants allowed rapid interpretation of high-likelihood pathogenic variants in relevant genes. The hands-on time for starting pipeline components and interpretation of known disease genes was, on average, less than 1 hour. Because genomic knowledge is currently limited to 1 to 2% of physicians (physician scientists, medical geneticists, and molecular pathologists), variant characterization, interpretation, and clinical guid- ance tools are greatly needed, as is training of medical geneticists and genetic counselors in their use. Return of results In blinded, retrospective analyses of two patients, rapid WGS correctly recapitulated known diagnoses. In child UDT002, two heterozygous, known mutations were identified in a gene that matched all clinical features. In male UDT173, a hemizygous (X-linked) VUS was identi- fied in the single candidate gene matching all clinical features. The variant, a nonsynonymous nucleotide substitution, was predicted to be damaging. Rapid WGS also provided a definitive diagnosis in one of four infants enrolled prospectively. In CMH172, with refractory epilepsy, rapid WGS disclosed a novel, homozygous frame-shifting insertion in a single candidate gene (BRATI). BRATI mutations were very recently reported in two unrelated Amish infants who suffered lethal, multifocal seizures (46). A molecular diagnosis was reached within 1 hour of WGS data inspection in CMH172, even though a- tant reference databases [Human Gene Mutation Database (HGMD) and OMIMI had not yet been updated with a BRATI disease associ- ation. The diagnosis was made clinically reportable by resequencing the patient and her parents. Had this diagnosis been obtained in real time, it may have expedited the decision to reduce or withdraw support. The latter decision was made in the absence of a molecular diagnosis after 5 weeks of ventilatory support, testing, and unsuc- itr cessful interventions to control seizures. Given high rates of NICU abed occupancy, accelerated diagnosis by rapid WGS has the potential to reduce the number of neonates who are turned away. The molec- ular diagnosis was also useful for genetic counseling of the infant's parents to share the information with other family members at risk 2 for carrying of this mutation. As suggested by recent guidelines (9), -,9 this case demonstrates the use of WGS for diagnostic testing when IL a genetic test for a specific gene of interest is not available. In four of five affected individuals, prospective, rapid WGS provided IF a definitive or likely molecular diagnosis in —50 hours. These cases dem- tel onstrated the use of WGS for diagnostic testing when a high degree of as genetic heterogeneity exists, as suggested by recent guidelines (9). Con- I firmatory resequencing, which is necessary for return of results until rapid c WGS is compliant with Clinical laboratory Improvement Amendments •—cDo (CLIA), took at least an additional 4 days. Until compliance has been established, we suggest preliminary verbal disclosure of molecular rt; diagnoses to the neonatologist of record, followed by formal reporting E upon performance of CLIA-conforming resequencing. Staged return of 2 results of broad or complex screening tests, together with considered, a- pert interpretation and targeted quantification and confirmation, is likely .8 to be acceptable in intensive care. Precedents for rapid return of interim, 8 potentially actionable results include preliminary reporting of histo- pathology, radiographic, and imaging studies and interim antibiotic selec- g tion based on Gram stains pending culture and sensitivity results. Disease gene sleuthing Because at least 3700 monogenic disease genes remain to be identified (10), WGS will often rule out known molecular diagnoses and suggest novel candidate disease genes (23, 51). Indeed, in another prospectively enrolled family, WGS resulted in the identification of a novel candidate disease gene, providing a likely molecular diagnosis. The proband was the second affected child of healthy parents. Accurate genetic counseling regarding risk of recurrence had not been possible because the first affected child lacked a molecular diagnosis. We undertook rapid WGS of the quartet simultaneously, allowing us to further limit incidental variants by requiring recessive inheritance. Rapid WGS ruled out 14 genes known to be associated with visceral heterotaxy and con- genital heart disease (HTX). Among genes that had not been associated with HTX, rapid WGS of the quartet narrowed the likely pathogenic variants to two in the BCL9L gene BCL9L had not previously been as- sociated with a human phenotype but is an excellent candidate gene for www.ScienceTranslationalMediclne.org 3 October 2012 Vol 4 Issue 54 154ral 3S 6 EFTA00315106 RESEARCH ARTICLE HTX based on its role in the Wingless (Writ) signaling pathway, which controls numerous developmental processes, including early embryonic patterning, epithelial-mesenchymal interactions, and stem cell mainte- nance (51, 52). Recently, the Writ pathway was implicated in the left-right asymmetric development of vertebrate embryos, with a role in the regulation of ciliated organ formation and function (53-57). The key effector of Writ signaling is P-catenin, which functions either to promote cell ad- hesion by linking cadherin to the actin cytoskeleton via a-catenin or to bind transcriptional coactivators in the nucleus to activate the ex- pression of specific genes (58-60). The protein that controls the switch between these two processes is encoded by BCL9L (also known as BCL9-2) and serves as a docking protein to link P-catenin with other transcription coactivators. BCL9L and a-catenin share competitive overlapping binding sites on (3-catenin; phosphorylation of P-catenin determines which pathway is activated. The mutation found in our patients lies within the BCL9L nuclear localization signal, which is essential for p•catenin to perform transcriptional regulatory functions in the nucleus (61). BCL9L is one of two human homologs of Drosophila legless (Igs), a segment polarity gene required for Wnt signaling during develop- ment lgs-deficient flies die as pharate adults with Wnt-related defects, induding absent legs, and antennae and occasional wing defects (62). Fly embryos lacking the maternal Igs contribution display a lethal seg- ment polarity defect. BCL9L-deficient zebrafish exhibit patterning de- fects of the ventrolateral mesoderm, including severe defects of trunk and tail development (60). Furthermore, inhibition of zebrafish $-catenin results in defective organ laterality (54). Overexpression of constitutively active P-catenin in medaka fish causes cardiac laterality defects (63). p•Catenin-deficient mice have defective development of heart, intes- tine, liver, pancreas, and stomach, including inverted cell types in the esophagus and posteriorization of the gut (64). Down-regulation of Wnt signaling in mouse and zebrafish causes randomized organ lat- erality and randomized side-specific gene expression. These likely re- flect aberrant Wnt activity on midline formation and function of Kupffer's vesicle, a ciliated organ of asymmetry in the zebrafish embryo that ini- tiates left-right development of the brain, heart, and gut (56, 65). The second human homolog of 1gs, BCL9, has been implicated in complex congenital heart disease in humans, of the type found in our patients (66-68). BCL9 was originally identified in precursor B cell acute lym- phoblastic leukemia with a t(1:14)(q21;q32) translocation (69), linking the Wnt pathway and certain B cell leukemias or lymphomas (62). Finally, it was recently demonstrated that the Wnt/P-catenin signaling pathway regulates the ciliogenic transcription factor foxjla expression in zebrafish (57). Decreased Wnt signal leads to disruption of left-right patterning, shorter/fewer cilia, loss of ciliary motility, and decreased foxjla expression. Foxjla is a member of the forkhead gene family and regulates transcriptional control of production of motile cilia (70). On the basis of this collected evidence, the symbol 1-11X6 has been reserved for BCL9L-associated autosomal recessive visceral heterotaxy. Additional studies are in progress to show causality definitively. These findings support clinical WGS as being valuable for research in reverse- translation studies (bedside to bench) that reveal new genetically ame- nable disease models. Addressing limitations In one remaining prospective patient, rapid WGS failed to yield a potential or definitive molecular diagnosis. Currently, WGS cannot survey every nucleotide in the genome (71). At 50x aligned coverage of the genome, WGS genotyped at least 95% of the reference genome with greater than 99.95% accuracy, using methods very similar to those used in this study (72). It has been suggested that this level of completeness is applicable for analyzing personal genomes in a clinical setting (72). In particular, GC-rich first exons of genes tend to be underrepresented (33). More complete clinical use of WGS will require higher sequencing depth, multi- platform sequencing and/or alignment methodologies, complementation by exome sequencing, or all three (73). Combined alignments with two methods of sequencing identified —9% more nucleotide variants than one alone. However, these additions raise the cost of WGS, increase the time to clinical interpretation, and shift the cost-benefit balance. For genetic disease diagnosis, the genomic regions that harbor known or likely rlits-Aci. mutations —the Mendelianome (2, 33)-must be genotyped accurately. In addition to exons and exon-intron bound- aries, the Mendelianome indudes some regions in the vicinity of genes that have structural variations or rearrangements. NGS of genome re- n gions that contain pseudogenes, paralogs (genes related by genomic duplication), or repetitive motifs can be problematic CMH064 had cc fulminant EB. Most EB-associated genes encode large cytoskektal pro- teins with regions of constrained amino acid usage, which equate with cI,; low nucleotide complexity. In addition, several EB-associated genes 2 have closely related paralogs or pseudogenes. These features impede unambiguous alignment of short reads, which can complicate attribu LL- tion of variants by NGS. This limitation can prevent definitive exdu- sion of candidate genes. For example, 45% of nucleotides in KRT14, rn an EB-associated gene, had <16-fold high-quality coverage and, thus, tel may have failed to disclose a heterozygous variant. In CMH064, how- al ever, this possibility was exduded by targeted sequencing of the re- gions of KRT14 known to contain mutations that cause EB. Furthermore, WGS is not yet effective for clinical-grade detection .2O of all mutation types. Copy number variations and large deletions 01 require clinical validation of research methods (33). Long, simple sequence-repeat expansions and complex rearrangements are prob- E lematic. Nevertheless, with WA-type adherence to standard operation- 2 al processes, the component of the Mendelianome for which WGS is effective is extremely reproducible (33). Thus, the specific disns.c, genes, exons, and mutation classes that are qualified for analysis, cc?, interpretation, and clinical reporting with WGS can be precisely pre- t dieted. This is of critical importance for reporting of differential diagnoses in the genetic disease arena Thus, although insufficient alone, rapid WGS may still be a cost-effective initial screening tool for differential diagnosis of EB. In our study, all EB-associated genes had >95% nudeotides with high-quality coverage sufficient to exclude heterozygous and homozygous nudeotide variants (>16-fold); 19 of these genes had >99% nucleotides with this coverage. Hence, for rig- orous testing of all EB-associated genes and mutation types, additional studies remain necessary, such as immunohistochemistry, targeted se- quencing of unca0able nucleotides, and cytogenetic studies. Of 531 disease genes examined, 52 had pseudogenes, paralogs, repetitive mo- tifs, or mutation types that may complicate WGS for comprehensive mutation detection. The comprehensiveness of WGS will be enhanced by longer reads, improved alignment methods, and validated algo- rithms for detecting large or complex variants (2, 4). Finally, in singleton (sporadic) cases, such as CMH064, family his- tory is often unrevealing in distinguishing the pattern of inheritance. For example, inheritance of ES can be dominant or recessive. Of two plausible heterozygous VUS detected in candidate genes in

📷 Images in this document (14 detected; 6 largest described)

AI-generated factual descriptions of embedded images (llava:13b). These are searchable across the corpus.

[Image 1] The image is a scanned document, specifically a research article from a scientific journal. The document is titled "Research Article" and is authored by "C.A. et al." The visible text includes the title "The role of the immune system in the pathogenesis of multiple sclerosis" and the names of the authors. There are also references to "Figure 1" and "Figure 2," which are likely to be images related [Image 2] The image appears to be a page from a scientific or medical journal. It contains text and two photographs. The text is too small to read clearly, but it seems to be related to a study or research article. The photographs show two different images: 1. The top photograph shows a close-up of a person's skin with visible lesions or rashes. The skin appears inflamed or irritated. 2. The bottom photogr [Image 3] The image is a page from a scientific or medical journal. It contains a list of steps or instructions, which are likely related to a research study or clinical trial. The text is in English and includes phrases such as "Obtain consent and blood sample," "Enter clinical findings into SASA," and "Perform sequencing library preparation." There are also references to "HGVS," "CASAVA," "SAGA," and "WGS [Image 4] The image shows a page from a scientific or medical journal. The page contains text, which appears to be an article or research paper. The text is dense and includes references to scientific studies, methods, and findings. There are no visible names, dates, places, or logos that can be discerned from this image. The content of the text is not described, as per the instructions. [Image 5] The image shows a page from a scientific or medical journal. The text is dense and appears to be discussing research related to genetics and cancer. There are references to genetic markers, mutations, and statistical analysis. The page is numbered, and there are headers and footers typical of academic papers. The text is too small to read in detail, but it seems to be a formal, scholarly publicati [Image 6] The image shows a page from a scientific or medical journal. The text is dense and appears to be discussing research related to a specific topic, possibly related to genetics or biology. The page is numbered and contains a header with the title "Research Article" followed by the name of the journal. There are sections with headings such as "Introduction," "Methods," "Results," and "Discussion," wh