Abstract
Major depressive disorder (MDD) diagnosis currently relies on subjective assessments, limiting early detection and accurate risk stratification. To address this, we developed machine learning models for both prevalent diagnosis and 15-year incident risk prediction using multimodal data from 30,040 UK Biobank participants. We integrated six modalities - proteomics, metabolomics, lifestyle factors, family history, biochemical measures, and polygenic risk scores (PRS) - employing a geographically stratified design with England-based participants for training and those from Scotland and Wales for independent validation. The integrated multimodal model demonstrated superior robustness, yielding a stable mean time-dependent AUC of 0.679 for incident prediction and an AUC of 0.791 for prevalent diagnosis. Crucially, our analysis revealed that modifiable lifestyle factors, particularly sleep patterns, and proteomic profiles consistently outperformed polygenic risk scores, which showed limited independent discriminative ability. Notably, the diagnostic model exhibited high specificity, effectively distinguishing MDD from schizophrenia (AUC = 0.865), bipolar disorder (AUC = 0.799), and anxiety disorders (AUC = 0.729). Feature attribution revealed distinct molecular signatures for disease stages: inflammatory mediators including CTSL and MMP7 dominated diagnosis, whereas markers of neuroplasticity and metabolism such as LRRN1 and LEP emerged as key predictors of future risk. Furthermore, biological enrichment analyses implicated dysregulated immune regulation (specifically neutrophil degranulation), lipid metabolism, and signal transduction pathways, highlighting a potential mechanistic shift from metabolic vulnerability in the risk phase to systemic inflammation in the symptomatic state. These findings demonstrate that integrating multimodal data significantly improves the characterization of MDD, offering a scalable framework for objective biomarker discovery and personalized risk assessment.</p>