user Admin_Adham
31st Jul, 2026 12:00 AM
Test

Machine Learning Model Boosts Genetic Prediction of T1D Risk

TOPLINE

A machine learning model, T1GRS, that used 160 genetic risk signals identified by researchers, improved prediction of type 1 diabetes (T1D) compared to a previous genetic risk score (GRS) in Europeans. The model’s improved prediction was particularly notable in people without high-risk human leukocyte antigen (HLA) haplotypes. Although the model was trained on variants identified in European ancestry, the model performed similarly to an ancestry-specific score in African Americans. In addition, T1GRS revealed four distinct genetic subtypes with significant differences in age of onset and rates of diabetes-related complications.

METHODOLOGY

  • T1D is largely inherited,driven mainly by class I and II HLA genes, but many other genes also contribute; larger studies and fine-mapping can identify more causal variants. Genetic risk scores generated from these findings can help predict who is at a risk, guide prevention strategies, and distinguish T1D from other forms of diabetes.
  • Researchers performed a genome-wide association analysis in 817,718 individuals of European ancestry (20,355 patients with T1D and 797,363 individuals without diabetes) and a fine-mapping analysis in 29,746 individuals (10,107 with T1D and 19,639 without) at the major histocompatibility complex (MHC) locus to identify genetic risk signals linked to T1D.
  • A machine learning model called T1GRS was trained using genetic variants at these signals, with validation in independent European and African American cohorts of patients with T1D, patients with T2D, and individuals without diabetes.
  • A clustering analysis was performed on 29,746 individuals based on features derived from T1GRS to identify distinct clinical characteristics associated with genetic subtypes, and these subtypes were replicated in an independent validation cohort.

TAKEAWAY

  • Overall, 160 risk signals at 97 risk loci plus the MHC locus were identified; the gradient-boosting machine learning model T1GRS was built using 199 variants (102 non-MHC and 97 MHC variants) and trained in two forms, one including genetic variants, sex, and ancestry principal components, and the other using only genetic variants. Separate MHC-only and non-MHC-only submodels were also developed.
  • T1GRS demonstrated improved classification of T1D compared with a conventional GRS model in 29,746 individuals of European ancestry (area under the receiver operating characteristic curve [AUC], 0.937 vs 0.916; DeLong test P = 5.31 × 10-21) and the validation cohort (AUC, 0.872 vs 0.791; DeLong test P = 1.14 × 10-16); the model performed comparably to current standards in predicting the risk for T1D in African Americans.
  • T1GRS reduced both missed and incorrect T1D classifications compared with the GRS models, especially in people who did not carry the common high-risk HLA haplotypes; the model also revealed 154 variant pairs with significant nonlinear interactions, indicating the model’s improved potential to predict the risk for T1D.
  • Individuals were categorized into four genetic subtype groups — MHC-driven, MHC-enriched, T cell-enriched, and pancreas-enriched — each marked by different key genes and cell-type signals. These groups showed different clinical patterns: The MHC-related subtypes developed T1D earlier, whereas the pancreas-enriched subtype had a later onset but had higher rates of renal, neurologic, and cardiac complications.

IN PRACTICE

“Our results highlight the value of combining the results of genetic association studies with machine learning methods to improve the prediction of complex diseases,” the authors wrote.

SOURCE

The study was led by Carolyn McGrail, Timothy J. Sears, and Emily N. Griffin, University of California, San Diego. It was published online in Nature Genetics.

LIMITATIONS

Accurate definition of T1D from electronic health records and laboratory measurements remains a challenge in population-based biobanks. As T1D is a complex disease with both genetic and environmental components, there are inherent limitations to the predictive ability of genetic data alone. The association study did not extensively consider rare risk variants, which may further improve prediction.

DISCLOSURES

This study received internal funding from the University of California, San Diego, through an Endowed Chair and support from the Mark Foundation for Cancer Research. One author disclosed serving as a consultant, receiving honoraria, and holding shares in certain pharma companies. Some other authors reported having a pending patent related to machine learning methods developed in this study.

SUGGESTED FOR YOU

This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.


Share This Article

Comments

Leave a comment