By harnessing advanced AI, MethylGPT decodes DNA methylation with unprecedented accuracy, offering new paths for age prediction, disease diagnosis, and personalized health interventions.
In a recent study posted to the bioRxiv preprint* server, researchers developed a transformer-based foundation model, MethylGPT, for the DNA methylome.
DNA methylation is a type of epigenetic modification that regulates gene expression via methyl-binding proteins and changes in chromatin accessibility. It also helps maintain genomic stability through transposable element repression. DNA methylation has features of an ideal biomarker, and studies have revealed distinct methylation signatures across pathological states, allowing for molecular diagnostics.
Nevertheless, several analytic challenges impede the implementation of diagnostics based on DNA methylation. Current approaches rely on simple statistical and linear models, which are limited in capturing complex, non-linear data. They also fail to account for context-specific effects such as higher-order interactions and regulatory networks. Therefore, a unified analytical framework that can model complex, non-linear patterns in various tissue and cell types is urgently needed.
Recent advances in foundation models and transformer architectures have revolutionized analyses of complex biological sequences. Foundation models have also been introduced for various omics layers, such as AlphaFold3 and ESM-3 for proteomics and Evo and Enformer for genomics. The achievements of the foundation models suggest that DNA methylation analyses could be transformed with a similar approach.
The study and findings
In the present study, researchers developed MethylGPT, a transformer-based foundation model for the DNA methylome. First, they acquired data on 226,555 human DNA methylation profiles spanning multiple tissue types from the EWAS Data Hub and Clockbase. Following deduplication and quality control, 154,063 samples were retained for pretraining. The model focused on 49,156 CpG sites, which were selected based on their known associations with various traits, as this would maximize their biological relevance.
The model was pre-trained using two complementary loss functions: masked language modeling (MLM) loss and profile reconstruction loss, enabling it to accurately predict methylation at masked CpG sites. The model achieved a mean squared error (MSE) of 0.014 and a Pearson correlation of 0.929 between predicted and actual methylation levels, indicating high predictive accuracy. Researchers also evaluated whether the model could capture biologically relevant features of DNA methylation. As such, they analyzed the learned representations of CpG sites in the embedding space.
They found that CpG sites clustered based on their genomic contexts, suggesting that the model learned the regulatory features of the methylome. In addition, there was a clear separation between autosomes and sex chromosomes, indicating that MethylGPT also captured higher-order chromosomal features. Next, the team analyzed zero-shot embedding spaces. This showed a clear biological organization, clustering by sex, tissue type, and genomic context.
Major tissue types formed well-defined clusters, indicating that the model learned methylation patterns specific to tissues without explicit supervision. Notably, MethylGPT also avoided batch effects, which often confound results in complex datasets. Besides, female and male samples demonstrated consistent separation, reflecting sex-specific differences. Next, the researchers assessed the ability of MethylGPT to predict chronological age from methylation patterns. To this end, they used a dataset of over 11,400 samples from diverse tissue types.
Fine-tuning for age prediction led to robust age-dependent clustering. Notably, intrinsic age-related organization was evident even before fine-tuning. Moreover, MethylGPT outperformed existing age prediction methods (e.g., Horvath’s clock and ElasticNet), achieving superior accuracy. Its median absolute error for age prediction was 4.45 years, further demonstrating its robustness. MethylGPT was also remarkably resilient to missing data. It exhibited stable performance with up to 70% missing data, outperforming multi-layer perceptron and ElasticNet approaches.
Analysis of methylation profiles during induced pluripotent stem cell (iPSC) reprogramming showed a clear rejuvenation trajectory; samples progressively transitioned to a younger methylation state over the course of reprogramming. The model was also able to identify the point during reprogramming (day 20) when cells began showing clear signs of epigenetic age reversal. Finally, the model’s ability to predict disease risk was assessed. The pre-trained model was fine-tuned to predict the risk of 60 diseases and mortality. The model achieved an area under the curve of 0.74 and 0.72 on validation and test sets, respectively.
In addition, they used this disease risk prediction framework to evaluate the impact of eight interventions on predicted disease incidence. Interventions included smoking cessation, high-intensity training, and the Mediterranean diet, among others, each of which showed varying degrees of effectiveness across disease categories. This showed distinct intervention-specific effects across disease categories, highlighting the potential of MethylGPT in predicting intervention-specific outcomes and optimizing tailored intervention strategies.
Conclusions
The findings illustrate that transformer architectures could effectively model DNA methylation patterns while preserving biological relevance. The organization of CpG sites based on regulatory features and genomic context suggests that the model captured fundamental aspects without explicit supervision. MethylGPT also demonstrated superior performance in age prediction across different tissues. Moreover, its robust performance in handling missing data (≤ 70%) underscores its potential utility in clinical and research applications.
News
X-chromosome study reveals hidden genetic links to Alzheimer’s disease
Despite decades of research, the X-chromosome’s impact on Alzheimer’s was largely ignored until now. Explore how seven newly discovered genetic loci could revolutionize our understanding of the disease. Conventional investigations of the genetic contributors [...]
The Unresolved Puzzle of Long COVID: 30% of Young People Still Suffer After Two Years
A UCL study found that 70% of young people with long Covid recovered within 24 months, but recovery was less likely among older teenagers, females, and those from deprived backgrounds. Researchers emphasized the need [...]
Needle-Free: New Nano-Vaccine Effective Against All COVID-19 Variants
A new nano-vaccine developed by TAU and the University of Lisbon offers a needle-free, room-temperature-storable solution against COVID-19, targeting all key variants effectively. Professor Ronit Satchi-Fainaro’s lab at Tel Aviv University’s Faculty of Medical and [...]
Photoacoustic PDA-ICG Nanoprobe for Detecting Senescent Cells in Cancer
A study in Scientific Reports evaluated a photoacoustic polydopamine-indocyanine green (PDA-ICG) nanoprobe for detecting senescent cells. Senescent cells play a role in tumor progression and therapeutic resistance, with potential adverse effects such as inflammation and tissue [...]
How Dysregulated Cell Signaling Causes Disease
Cell signaling is crucial for cells to communicate and function correctly. Disruptions in these pathways, caused by genetic mutations or environmental factors, can lead to uncontrolled cell growth, improper immune responses, or errors in [...]
Scientists Develop Super-Strong, Eco-Friendly Plastic That Bacteria Can Eat
Researchers at the Weizmann Institute have developed a biodegradable composite material that could play a significant role in addressing the global plastic waste crisis. Billions of tons of plastic waste clutter our planet. Most [...]
Building a “Google Maps” for Biology: Human Cell Atlas Revolutionizes Medicine
New research from the Human Cell Atlas offers insights into cell development, disease mechanisms, and genetic influences, enhancing our understanding of human biology and health. The Human Cell Atlas (HCA) consortium has made significant [...]
Bioeconomic Potential: Scientists Just Found 140 Reasons to Love Spider Venom
Researchers at the LOEWE Centre for Translational Biodiversity Genomics (TBG) have discovered a significant diversity of enzymes in spider venom, previously overshadowed by the focus on neurotoxins. These enzymes, found across 140 different families, [...]
Quantum Algorithms and the Future of Precision Medicine
Precision medicine is reshaping healthcare by tailoring treatments to individual patients based on their unique genetic, environmental, and lifestyle factors. At the forefront of this revolution, the integration of quantum computing and machine learning [...]
Scientists Have Discovered a Simple Supplement That Causes Prostate Cancer Cells To Self-Destruct
Menadione, a vitamin K precursor, shows promise in slowing prostate cancer in mice by disrupting cancer cell survival processes, with potential applications for human treatment and myotubular myopathy therapy. Prostate cancer is a quiet [...]
Scientists reveal structural link for initiation of protein synthesis in bacteria
Within a cell, DNA carries the genetic code for building proteins. To build proteins, the cell makes a copy of DNA, called mRNA. Then, another molecule called a ribosome reads the mRNA, translating it [...]
Vaping Isn’t Safe: Scientists Uncover Alarming Vascular Risks
Smoking and vaping impair vascular function, even without nicotine, with the most significant effects seen in nicotine-containing e-cigarettes. Researchers recommend avoiding both for better health. Researchers have discovered immediate impacts of cigarette and e-cigarette [...]
Twice-Yearly Lenacapavir for PrEP Reduces HIV Infections by 96%
Twice-yearly injections of the capsid inhibitor drug lenacapavir can prevent the vast majority of HIV infections, according to a Phase 3 clinical trial published Wednesday in the New England Journal of Medicine. HIV pre-exposure [...]
Did Social Distancing Begin 6,000 Years Ago? Neolithic Villagers May Have Invented It
Social distancing may have roots 6,000 years ago, as research shows Neolithic villages like Nebelivka used clustered layouts to control disease spread. The phrase “social distancing” became widely recognized in recent years as people [...]
Decoding Alzheimer’s: The Arctic Mutation’s Role in Unusual Brain Structures
Researchers have uncovered how certain genetic mutations lead to unique spherical amyloid plaques in inherited forms of Alzheimer’s, offering insights that could advance our understanding of the disease and improve therapeutic strategies. An international collaboration [...]
How Your “Lizard Brain” Fuels Overthinking and Social Anxiety
New research by Northwestern Medicine reveals how humans have evolved advanced brain regions to interpret others’ thoughts, connecting these areas with the amygdala, a part of the brain involved in emotional processing. Study sought [...]