About
This project aims to leverage genome sequence data from the UK Biobank to improve the prediction of complex traits and deepen our understanding of the genetic architecture underlying human phenotypic variation. While genome-wide association studies (GWAS) using imputed genotype arrays have identified many common variants associated with complex traits, whole-exome sequence (WES) and whole-genome sequence (WGS) data enable a more comprehensive capture of both common and rare variants, including those in poorly imputed or non-coding regions.
Research questions include:
(1) To what extent can WGS data improve the prediction accuracy of complex traits compared to array-based or imputed genotypes?
(2) What is the contribution of low-frequency and rare variants to trait heritability and prediction?
(3) How do functionally annotated variants contribute to phenotypic variance, and what biological insights can be gained by integrating functional annotations?
The objectives are to:
* Develop and apply statistical models to predict complex traits (e.g., height, blood pressure, type 2 diabetes) using WES/WGS data.
* Quantify the relative contributions of common vs. rare variants to trait prediction and heritability.
* Integrate functional annotations, including those derived from transcriptomic, epigenomic, and proteomic datasets at bulk-tissue or single-cell level, into prediction and fine-mapping models to improve interpretability and biological inference.
Scientific rationale:
WES/WGS data provides unprecedented resolution to study the full spectrum of genetic variation. By harnessing this richness and integrating functional genomic data, we aim to not only enhance predictive models but also identify biologically meaningful variants and cellular contexts that inform disease mechanisms. This integrative approach will support the long-term goals of precision medicine by improving both risk stratification and understanding of trait biology.