Title: | PERHAPS: Paired-End short Reads-based HAPlotyping from next-generation Sequencing data |
Journal: | Briefings in Bioinformatics |
Published: | 8 Dec 2020 |
Pubmed: | https://pubmed.ncbi.nlm.nih.gov/33285565/ |
DOI: | https://doi.org/10.1093/bib/bbaa320 |
Title: | PERHAPS: Paired-End short Reads-based HAPlotyping from next-generation Sequencing data |
Journal: | Briefings in Bioinformatics |
Published: | 8 Dec 2020 |
Pubmed: | https://pubmed.ncbi.nlm.nih.gov/33285565/ |
DOI: | https://doi.org/10.1093/bib/bbaa320 |
WARNING: the interactive features of this website use CSS3, which your browser does not support. To use the full features of this website, please update your browser.
The identification of rare haplotypes may greatly expand our knowledge in the genetic architecture of both complex and monogenic traits. To this aim, we developed PERHAPS (Paired-End short Reads-based HAPlotyping from next-generation Sequencing data), a new and simple approach to directly call haplotypes from short-read, paired-end Next Generation Sequencing (NGS) data. To benchmark this method, we considered the APOE classic polymorphism (*1/*2/*3/*4), since it represents one of the best examples of functional polymorphism arising from the haplotype combination of two Single Nucleotide Polymorphisms (SNPs). We leveraged the big Whole Exome Sequencing (WES) and SNP-array data obtained from the multi-ethnic UK BioBank (UKBB, N=48,855). By applying PERHAPS, based on piecing together the paired-end reads according to their FASTQ-labels, we extracted the haplotype data, along with their frequencies and the individual diplotype. Concordance rates between WES directly called diplotypes and the ones generated through statistical pre-phasing and imputation of SNP-array data are extremely high (>99%), either when stratifying the sample by SNP-array genotyping batch or self-reported ethnic group. Hardy-Weinberg Equilibrium tests and the comparison of obtained haplotype frequencies with the ones available from the 1000 Genome Project further supported the reliability of PERHAPS. Notably, we were able to determine the existence of the rare APOE*1 haplotype in two unrelated African subjects from UKBB, supporting its presence at appreciable frequency (approximatively 0.5%) in the African Yoruba population. Despite acknowledging some technical shortcomings, PERHAPS represents a novel and simple approach that will partly overcome the limitations in direct haplotype calling from short read-based sequencing.</p>
Application ID | Title |
---|---|
19416 | A genome-wide and Pheno-wide association study of common diseases on 900,000 individuals from US and UK. |
Enabling scientific discoveries that improve human health