Document text
Principal Investigator: FREDERICK R BLATTNER
Organization: DNASTAR, INC.
Fiscal Year: 2024
Award: $681,914
Funding agency: National Institute of General Medical Sciences
The B cell population in each individual produces an estimated 1010 different antibodies, collectively known as
the antibody repertoire. This extraordinary diversity is essential for responding to the unique history of
infections, vaccinations and cancer encountered over an individual’s lifetime. Conversely, regulatory errors in
the system play a pivotal role in a host of auto-immune diseases. Antibodies are composed of two proteins, a
heavy and light chain, each containing a variable region, VH and VL, which together confer antigen binding
specificity. Diversity is initiated through differential recombination at the three V region encoding loci to produce
the naïve repertoire. Upon antigen exposure, B cells expressing an antibody specific to that antigen undergo
clonal expansion and concentrated somatic hypermutation (SHM) of V region sequences that code for the
antigen recognition domain. Those clonally derived B cells (clonotypes) each express a different sequence and
thereby structural variant of the initial unmutated antibody. Cells expressing higher affinity variants are selected
for in a process known as affinity maturation. In this way, the mature repertoire is built from the history of
antigenic encounters by that individual. Efficient deciphering of that history could contribute to improving
human health in numerous ways from better clinical decision making to improved diagnostics and therapeutics.
Toward that goal, ongoing technological advances in both DNA/RNA sequencing, protein structure modeling
software and high-performance scalable computer hardware are making virtual repertoire scale antibody
structure and antigen screening attainable in the not-too-distant future.
In this Direct to Phase II application, we propose to build a software suite that bridges the gap between
genomics and structural biology enabling antibody repertoires to be deciphered and mined in exquisite detail.
To do so, we first leverage our highly extensible sequence assembler, XNG, to produce haplotype phased and
annotated sequences of the germline IG loci from which the naïve repertoire can be simulated (Aim 1). Next,
XNG is used to assemble and annotate bulk BCR-seq data producing the linear VH and VL encoding
sequences of the mature repertoire (Aim 2). Translated repertoire sequences are then used as input for our
protein modeling software, NovaFold-Ab and NovaFold-AI, where high accuracy 3D antibody structures are
predicted (Aim 3). Those antibody structure libraries are then used in virtual screens to identify members that
bind to a target antigen with our protein interaction modeling program, NovaDock (Aim 4). Screens can also be
refined to specific epitopes of interest, for example, those known to elicit neutralizing antibodies. If realized,
these capabilities will have significant commercial opportunities for complementing existing technology in
improving clinical care and personalized medicine as well as aiding in the development of faster, more cost
effective diagnostics and therapeutics.
Terms: <2019-nCoV S protein><2019-nCoV spike glycoprotein><2019-nCoV spike protein><3-D><3-D structure><3-Dimensional><3-dimensional structure><3D><3D structure><Ab response><Acceleration><Adaptive Immune System><Affinity><Allergens><Antibodies><Antibody Affinity><Antibody Formation><Antibody Production><Antibody Repertoire><Antigen Targeting><Antigenic Determinants><Antigens><Autoimmune Diseases><B blood cells><B cell><B cells><B-Cells><B-Lymphocytes><B-cell><B-cell receptor repertoire sequencing><B-cell receptor sequencing><BCR repertoire sequencing><BCR seq><BCR sequencing><BCRseq><Binding><Binding Determinants><Biology><Biotech><Biotechnology><COVID-19 S protein><COVID-19 spike><COVID-19 spike glycoprotein><COVID-19 spike protein><Cancers><Cell Body><Cells><Characteristics><Clonal Expansion><Code><Coding System><Communities><Complex><Computer Hardware><Computer software><Computers><DNA><DNA Recombination><Data><Deoxyribonucleic Acid><Development><Development and Research><Diagnostic><Distant><Emergent Technologies><Emerging Technologies><Epitopes><Future><Genes><Genetic><Genetic Recombination><Genomics><Germ Lines><Goals><Grouping><Haplotypes><Health><Healthcare><History><Hour><Human><Ig Somatic Hypermutation><Immune system><Immunoglobulin Somatic Hypermutation><Individual><Infection><Libraries><Light><Malignant Neoplasms><Malignant Tumor><Messenger RNA><Modeling><Modern Man><Molecular Configuration><Molecular Conformation><Molecular Immunology><Molecular Interaction><Molecular Stereochemistry><Nature><Performance><Phase><Photoradiation><Play><Population><Process><Proteins><R & D><R&D><RNA Seq><RNA sequencing><RNAseq><Recombination><Recording of previous events><Research><Research Resources><Resources><Role><Running><SARS-CoV-2 S><SARS-CoV-2 S protein><SARS-CoV-2 spike><SARS-CoV-2 spike glycoprotein><SARS-CoV-2 spike protein><Severe acute respiratory syndrome coronavirus 2 S protein><Severe acute respiratory syndrome coronavirus 2 spike glycoprotein><Severe acute respiratory syndrome coronavirus 2 spike protein><Software><Specialist><Specificity><Structure><System><Technology><Therapeutic><Translating><Vaccination><Validation><Variant><Variation><acquired immune system><antibody biosynthesis><antigen antibody affinity><antigen binding><antigen bound><autoimmune condition><autoimmune disorder><autoimmunity disease><clinical care><clinical decision-making><computer system hardware><computing hardware><conformation><conformational><conformational state><conformationally><conformations><contig><coronavirus disease 2019 S protein><coronavirus disease 2019 spike glycoprotein><coronavirus disease 2019 spike protein><cost effective><deep learning><deep learning method><deep learning strategy><develop software><developing computer software><developmental><diagnostic development><environmental agent><groupings><health care><histories><immunogen><immunoglobulin biosynthesis><improved><in silico><interest><mRNA><malignancy><member><neoplasm/cancer><neutralizing antibody><new diagnostics><next generation diagnostics><novel diagnostics><pathogen><personalization of treatment><personalized medicine><personalized therapy><personalized treatment><programs><protein purification><protein structure><protein structures><proteins structure><research and development><response><screening><screenings><social role><software development><somatic hypermutation><spike proteins on SARS-CoV-2><structural biology><therapeutic agent development><therapeutic development><three dimensional><three dimensional structure><tool><transcriptome sequencing><transcriptomic sequencing><validations><virtual>