| Title: | Causal inference for multiple risk factors and diseases from genomics data |
| Journal: | Nature Communications |
| Published: | 25 Aug 2026 |
| DOI: | https://doi.org/10.1038/s41467-026-76877-7 |
| URL: | http://dx.doi.org/10.1038/s41467-026-76877-7 |
| Title: | Causal inference for multiple risk factors and diseases from genomics data |
| Journal: | Nature Communications |
| Published: | 25 Aug 2026 |
| DOI: | https://doi.org/10.1038/s41467-026-76877-7 |
| URL: | http://dx.doi.org/10.1038/s41467-026-76877-7 |
WARNING: the interactive features of this website use CSS3, which your browser does not support. To use the full features of this website, please update your browser.
Mendelian randomization, the standard instrument variable causal-inference approach for genomics, can be severely biased by inadequate selection of appropriate instrument variables and horizontal pleiotropy. We introduce Causal Inference GWAS, a graphical approach that selects genetic instruments from summary statistics controlling for linkage and pleiotropy, accommodates rare variants and binary outcomes, distinguishes direct from indirect risk factors in high dimensions, and flags latent confounding. Applied to 9 risk factors, 4 metabolic disease outcomes, and 8.4M variants across 458,747 UK Biobank individuals, our method runs in 20 minutes, identifying only 696 genome-wide valid instruments. We replicate nearly all previously reported risk factor-outcome paths, but find that most are indistinguishable from unmeasured confounding. This suggests robust biobank causal inference will require longitudinal and family data to resolve temporal precedence from reverse causation. Our approach offers a principled first step toward screening genuine causal signal from phenotypic correlation in biobank data.</p>
| Application ID | Title |
|---|---|
| 35520 | Improving estimation and prediction of common complex disease risk |
Enabling scientific discoveries that improve human health