Models and Methods for Population Genomics

NIH Pandemic-Era Grants

Pandemic Era Grants

2024

Document text

Principal Investigator: JOHN D STOREY
Organization: PRINCETON UNIVERSITY
Fiscal Year: 2024
Award: $365,761
Funding agency: National Human Genome Research Institute

Project Summary
Title:
Models and Methods for Population Genomics
Abstract:
Understanding genome-wide genetic variation and its role in health-related complex traits in humans is one of
the most important goals of modern biomedical research. There continues to be a substantial need for new
statistical models and methods that can be applied in these studies, particularly as study designs become more
ambitious and sample sizes increase. The overarching goal of this grant is to develop statistical theory, methods,
and software useful in understanding population genomics studies that involve genome-wide genotyping, a wide
range of measured traits, very large sample sizes, structured populations, and varying study designs.
One of the most challenges aspects of modern population genomics studies is that there is a complex
evolutionary history underlying the present-day genetic variation that we observe. Individuals are members of
structured populations with varying levels of relatedness that do not follow the simple assumptions that underlie
classical population genetics theory. There is a need to model and estimate arbitrary forms of structure and
relatedness so that genetic variation in human populations can be accurately characterized, which in turn allows
for an accurate understanding of the genetic basis of complex traits. Our first focus is on flexible, broadly
applicable models that adapt to this arbitrary population structure and relatedness, resulting in principled
statistical methods that make accurate inferences. We then show how our methods improve the ability to identify
genetic associations, estimate genome-wide heritability of traits, and contribute to an understanding of how
predictive polygenic risk scores can be robustly constructed.
The specific aims involve (1) introducing a parametric framework for estimating kinship and FST, thereby bridging
identity-by-descent models with random allele frequency coancestry models of structure; (2) advancing models
and methods for quantifying genome-wide heritability, testing for associations, and building polygenic risk scores
by incorporating our new estimation framework of kinship and FST; (3) developing and distributing software; and
(4) analyzing important data sets to discover new biology and validate our methods and software.

Terms: <Admixture><Allele Frequency><Bio-Informatics><Bioinformatics><Biology><Biomedical Research><Complex><Computer software><Data Set><Disease><Disorder><GWA study><GWAS><Gene Frequency><Genetic><Genetic Differentiation><Genetic Divergence><Genetic Diversity><Genetic Drift><Genetic Variation><Genomics><Genotype><Goals><Grant><Health><Heritability><History><Human><IBD analysis><IBD inference><Individual><Linkage Disequilibrium><Measures><Medical Research><Methodology><Methods><Modeling><Modern Man><Modernization><Pattern><Performance><Polygenic Characters><Polygenic Inheritances><Polygenic Traits><Population><Population Genetics><Probabilistic Models><Probability Models><Programming Languages><R programming language><Recording of previous events><Research><Research Design><Role><Sample Size><Software><Source Code><Statistical Methods><Statistical Models><Structure><Study Type><Testing><Update><Writing><allelic frequency><flexibility><flexible><genetic association><genome scale><genome wide association><genome wide association scan><genome wide association studies><genome wide association study><genome-wide><genomewide><genomewide association scan><genomewide association studies><genomewide association study><histories><human disease><identical by descent><identity by descent><improved><member><polygenic risk score><semiparametric><social role><statistic methods><statistical linear mixed models><statistical linear models><study design><theories><trait><web site><website><whole genome association analysis><whole genome association studies><whole genome association study>