Epigenetic Aging Clocks as Longitudinal Health Biomarkers
Tracking biological age changes over time predicts mortality better than single snapshots.

Chronological age counts the years someone has lived. Biological age tries to capture something harder to pin down: how much wear and tear the body has actually accumulated, at the molecular level, regardless of what the birth certificate says. Epigenetic clocks are the main tool for making that estimate, and for years, most of the research and consumer attention around them has focused on a single question: what's your biological age right now, at this one snapshot in time? That framing turns out to be the least interesting use of the technology. The evidence accumulating over the last two years points somewhere else entirely: the real predictive power is in how a person's clock reading changes over months and years, not in where it is on any given day.
The mechanism behind these clocks is DNA methylation, a chemical mark that attaches to specific spots on the genome called CpG islands (short stretches of DNA rich in cytosine and guanine bases). Methylation patterns shift in predictable ways as cells age, and machine-learning models trained on those patterns can estimate biological age from a blood or saliva sample. A single reading tells you where someone sits relative to population norms at that moment. Most people who encounter one number assume that higher means worse and lower means better, full stop.
That assumption is incomplete. That assumption is incomplete rather than wrong. A baseline reading conflates two things that matter for very different clinical reasons: how much change has already piled up, and how fast new change is arriving. Those are not the same signal. A person whose biology looks ten years older than their driver's license might be stable, or might be in the middle of a steep decline. A single measurement cannot tell the difference. Tracking the same person's clock over time can, and that distinction is what the rest of this piece works through.
How clock design shapes what they can and cannot detect longitudinally
Not every epigenetic clock was built to do the same job, and that matters enormously once you start measuring someone twice instead of once.
First-generation clocks, Horvath's being the best known example, were trained on one task: predict chronological age from methylation data as accurately as possible. They're good at that. But because they were never trained against health outcomes, they tend to sit flat even when a person's underlying biology is shifting in ways that matter. Asking a first-generation clock to detect an intervention's effect is a bit like asking a bathroom scale to measure blood pressure. It's simply not built for that.
Second-generation clocks changed the training objective. PhenoAge, GrimAge, and GrimAge2 were built using multiple biomarkers, with GrimAge folding in smoking pack-years as well, and the goal shifted from estimating age to predicting mortality and disease risk directly. Third-generation clocks, DunedinPACE and DunedinPoAm among them, went further still: instead of estimating a static age, they estimate the current pace of aging, a rate rather than a point. That architecture makes them naturally suited to repeated measurement. Fourth-generation clocks, including CausAge, DamAge, and AdaptAge, use Mendelian randomization to select methylation sites believed to play a causal role in aging rather than just a correlational one. CausAge functions as a general-purpose clock, DamAge tracks accumulated age-related damage specifically, and AdaptAge tracks adaptive or protective changes, a distinction that matters when the goal is to see whether an intervention is helping or hurting.
Across the current literature, a fairly wide set of named clocks, bAge, CausAge, CheekAge, DNAm PhenoAge, DunedinPACE, DunedinPoAm, FitAge, GrimAge, GrimAge2, InflammAge, OMICmAge, and Systems Age, have each shown significant, independent associations with all-cause mortality in longitudinal cohorts. That list spans several generations, but the pattern within it is consistent: clocks trained on mortality or aging pace, not on chronological age alone, carry the signal.
It's a training-objective problem baked into the model from the start. It's a training-objective problem baked into the model from the start. Pick a clock built to estimate age, and use it to track whether a diet change worked, and you may see nothing move at all, because the tool was never built to notice, even if biology did change.
What longitudinal tracking adds that a baseline cannot provide
A single clock reading answers one question: how old does this person's biology look today? A trajectory answers a different one entirely: how fast is that number moving, and in which direction? Those two questions produce different clinical information, and conflating them is where a lot of the confusion around biological age testing comes from.
Consider two people with an identical baseline reading, both clocking in five years "older" than their chronological age. One of them has been stable at that gap for a decade. The other arrived there after two years of rapid acceleration. Nothing about the baseline number distinguishes them. Only repeated measurement can.
That's the core argument for why rate of change functions as an independent variable, not just a derivative of the starting point. Research reviewed later in this piece shows trajectory predicting mortality risk even after controlling for baseline age, which is the technical way of saying: knowing where someone started doesn't tell you what the slope of their line looks like, and the slope carries its own weight.
Direction matters as much as magnitude. A clock that has been accelerating over two years implies something different from one that's decelerated, even if the two end up at the same absolute value on the day of comparison. Think of it the way a cardiologist thinks about cholesterol. A single cholesterol reading tells you where a patient stands. A trend line across five years tells you whether their risk is being managed or getting away from them, and that's the information that actually changes what a doctor does next. Biological age works the same way. The trajectory is the signal. The baseline is just the coordinate where the line starts.
The InCHIANTI cohort evidence: longitudinal clock change independently predicts mortality
The clearest empirical demonstration of that principle so far comes from Kuo and colleagues, published in Nature Aging, using data from the InCHIANTI study, a cohort of 699 adults followed for up to 24 years.
The central finding: longitudinal changes in several epigenetic clocks predicted long-term mortality independent of baseline epigenetic age and other confounding factors. People whose clocks accelerated faster over the follow-up period faced higher mortality risk, and that relationship held even after accounting for where their clock started.
That "independent of baseline" qualifier does a lot of work. It rules out the simpler, less interesting explanation, that people who start out biologically older just die sooner. Instead, it says the rate of movement itself carries information the starting point does not. Four clocks showed this pattern clearly: DNAmGrimAge v2, DNAmGrimAge, DNAmPhenoAge, and DunedinPACE. Combined models incorporating both baseline reading and rate of change reached C-statistics between 0.800 and 0.808, a level of discrimination for all-cause mortality that's genuinely strong for an elderly cohort followed this long.
Twenty-four years of follow-up is unusual in this field, where most studies work with two time points spaced a few years apart. That duration gives the mortality associations a kind of temporal weight shorter studies simply can't match. What InCHIANTI does not resolve is whether deliberately slowing someone's clock trajectory through intervention produces the same mortality benefit observed here. That's a separate question, and one the intervention trial literature is only beginning to answer.
Cross-clock validation at scale: what 14 clocks across 18,859 adults established
If InCHIANTI proved the principle in a smaller, long-running cohort, Mavrommatis and colleagues, using the Generation Scotland cohort in 2025, tested whether it generalizes.
The scale here is the point: 14 clocks tested against 174 disease outcomes and all-cause mortality across 18,859 adults, the largest head-to-head comparison of its kind reported to date. The finding: second- and third-generation clocks received their strongest empirical validation yet, with respiratory and liver outcomes emerging as the domains where prediction was strongest.
Predictive power isn't uniform across disease categories, and this affects how much confidence clinicians can place in a given clock's readings for a specific disease. Predictive power isn't uniform across disease categories. Some organ systems apparently leave a clearer methylation signature of aging than others, at least given current clock designs, which suggests future clocks may need to specialize by organ system rather than aim for a single whole-body number; the organ-specific clock research addresses this directly later on.
Generation Scotland extends InCHIANTI's finding from a tightly studied cohort of 699 to a population nearly 27 times larger, across a far wider set of disease outcomes. What it doesn't resolve is a practical question that matters for actually using these clocks in clinics: most large studies still rely on a limited number of time points. How many measurements, spaced how far apart, are needed to reliably estimate an individual's trajectory rather than catching a temporary blip remains an open practical question.
Lifestyle and environment shift the trajectory: what modifiable factors do to clock rate
Research into which everyday factors move the needle on clock rate has produced a list that reads close to a standard cardiovascular risk profile.
Cardiometabolic risk factors have associated with faster pace-of-aging readings on DunedinPACE and related tools. Healthier lifestyle behaviors have pulled in the other direction, associated with slower pace readings. None of that will surprise anyone who follows cardiometabolic research, but the fact that it appears this clearly in methylation data is itself notable: these clocks appear to be picking up the same risk signal cardiologists have tracked for decades, just through a different molecular window.
Research has raised a layer of nuance: the relative weight of these factors may differ by sex, with the same modifiable risk factors potentially carrying different influence depending on who's being measured, which argues against treating any single lifestyle intervention as universally equivalent across a population.
People who shift into faster aging trajectories tend to show worsening cardiovascular profiles over time. Whether faster aging causes the cardiovascular decline, or the reverse, or both feed each other, remains genuinely open. What's clear is the association's shape, moving in step, not which side is pulling the other.
One caution belongs here before moving to intervention trials. These are population-level patterns. An individual who starts exercising more or brings their blood glucose down isn't guaranteed a proportional shift in their own clock reading. Population averages describe tendencies, not promises, and the intervention research below is where that gap gets tested more rigorously, person by person.
What intervention trials reveal about which clocks move
A 2026 Nature Medicine analysis pulled together evidence on 16 different epigenetic clocks across 51 longitudinal intervention studies, spanning both drug trials and lifestyle programs, and the pattern that emerged across all 51 is remarkably consistent: clocks trained to predict mortality or aging pace respond to intervention. Clocks trained only to predict chronological age, Horvath in particular, barely budge, for the same design reason discussed earlier. It was never built to notice.
Within the responsive clocks, there's a further split. Within the responsive clocks, the clock a trial picks determines what that trial is even capable of detecting, independent of what's actually happening in participants' bodies, a practical consequence that affects how intervention results should be interpreted.
The CALERIE trial, a randomized controlled study of caloric restriction in healthy adults, illustrates this precisely. DunedinPACE showed a slower pace of aging in the caloric restriction group. PhenoAge and GrimAge showed nothing. The clock-specific result underlines how much clock choice shapes what a trial can detect, independent of what's actually happening in participants' bodies.
A 12-week multimodal lifestyle trial from 2026, combining exercise, dietary guidance, and daily yogurt containing the probiotic strain Bifidobacterium longum BB536 in overweight men aged 50 and older, produced a significant DunedinPACE deceleration, translating to roughly a 2.2% slower pace of aging in the intervention group, with no meaningful shift in controls.
Supplementation research adds more texture. Omega-3 supplementation produced significant standardized estimates of -0.17 for DunedinPACE, -0.32 for GrimAge, and -0.16 for PC DNAm PhenoAge, translating to -0.0022 units per year on DunedinPACE, -0.64 years on GrimAge2, and -0.24 years on PC DNAm PhenoAge. Stacking omega-3 with vitamin D and exercise together produced an additive drop of -0.32 years on PC DNAm PhenoAge, suggesting these interventions may compound rather than simply overlap.
Metformin, tested in patients with diabetes, produced a reduction in epigenetic age acceleration of 2.77 years on the Horvath clock and 3.43 years on the Hannum clock. Horvath is the clock this section just described as largely unresponsive to intervention, which makes this shift notable. A first-generation clock moving this much under metformin likely reflects the drug's broad metabolic effects on methylation patterns generally, rather than a clean, isolated signal about aging rate specifically. It's a result that deserves attention without over-interpretation.
Put together, these trials confirm that clock trajectories can be shifted by real interventions. But the size of the shift, and even whether a shift occurs at all, depends heavily on which clock is measuring it, how long the intervention runs, and what kind of intervention it is. Not every positive result in this literature means the same thing.
Disease prediction beyond mortality: stroke, frailty, cancer, and organ-specific clocks
Mortality prediction is the headline use case, but it's far from the only one. A systematic review and meta-analysis synthesizing 13 studies found that accelerated biological aging consistently associated with higher stroke risk, and the association was stronger for first-ever stroke than for recurrent events. That pattern suggests these clocks may be picking up on vascular vulnerability that standard risk factor models don't fully capture on their own.
Frailty shows a related pattern: higher GrimAge epigenetic age acceleration consistently tracks with greater frailty. But the field openly acknowledges a gap here, calling for clocks specifically trained and validated against frailty outcomes in large longitudinal cohorts, rather than repurposing mortality-trained clocks for a related but distinct question.
Cancer research adds another dimension. A study in eBioMedicine used Mendelian randomization and methylation quantitative trait locus analysis to identify 15 aging-linked CpG sites whose methylation status appears to influence colorectal cancer risk, with genes including TNF, NCF2, BICC1, DIP2B, TBX3, and SUCNR1 implicated. Notably, accelerated aging showed stronger predictive value for early-onset colorectal cancer specifically, which hints that epigenetic drift might create vulnerability earlier in life than most conventional cancer risk models currently assume.
Perhaps the most structurally interesting development is the move toward organ-specific clocks. A Nature Aging study analyzed 43,616 UK Biobank participants using Olink's 3,072-protein measurement platform to build separate aging clocks for multiple organs. These organ-specific clocks strongly predicted disease risk and mortality tied to their respective systems. A separate Nature Medicine paper found that accelerated brain and immune system aging, measured through plasma proteins, were the strongest mortality and disease predictors among all the organ-specific clocks studied.
This changes the picture meaningfully. A single composite biological age number assumes the whole body ages as one unit. Organ-specific clocks show that's often not true: a person might be aging quickly in their liver while their cardiovascular system holds steady, or vice versa, and a whole-body average would simply hide that. For anyone using these tools to guide actual health decisions, that's not a minor technical footnote, it changes what the number is even supposed to mean.
The reliability problem that longitudinal use makes harder to ignore
None of the above matters if the measurement itself is noisy, and this is where longitudinal use runs into its hardest constraint. Interpreting change over time requires confidence that the change reflects real biology and not just measurement variability, and reliability differs a lot from clock to clock.
A pooled intraclass correlation (ICC) analysis across four independent studies, covering 18 commonly used clocks, found that most achieved excellent technical reliability, with ICC values above 0.9. SystemsAge, GrimAge, and PC-based versions of established clocks ranked consistently among the most stable performers. Earlier-generation clocks, Hannum, PhenoAge, and DNAm PhenoAge among them, showed somewhat lower but still generally acceptable reliability, in the "good" range around 0.7 to 0.8 ICC.
The bigger complication occurs under real-world conditions rather than controlled lab repeats. Short-term perturbations, a heavy meal, a stressful week, exposure to air pollution, caused substantial fluctuation in epigenetic age estimates across repeated measures, dropping most clocks down to only moderate reliability at best when tested this way. That has a direct design consequence: a reading taken right after a rough week at work, or an unusual stretch of eating, might not reflect someone's actual biological trajectory at all. Standardized sampling conditions matter far more in longitudinal designs than they do when a clock is only ever measured once.
A GeroScience paper by Hamaya and colleagues, published in November 2025, focused specifically on longitudinal changes in epigenetic measures over a two-year window and their methodological implications. The fact that a paper like this exists at all signals something about where the field's attention has shifted: reliability, not predictive theory, is now the binding constraint on whether longitudinal clock data can be trusted for anything clinical.
Principal-component versions of established clocks offer a partial fix, showing consistently higher technical reproducibility than their original counterparts. That's a real improvement, though it comes with a tradeoff: PC-based clocks generally sacrifice some biological specificity in exchange for that stability. What the field still lacks is any consensus on the minimum ICC a clock needs to clear before it's trustworthy for monitoring a single individual over time, as opposed to studying trends across a large population, a distinction that matters enormously and hasn't been settled.
What needs to be true before longitudinal clock data can guide clinical decisions
Several things have to fall into place before a doctor can look at a patient's clock trajectory and act on it the way they'd act on a rising blood sugar marker or a climbing cholesterol marker.
Measurement conditions need standardizing first. Sampling protocols have to account for short-term perturbations, meals, stress, acute exposures, so that a fluctuation gets separated from a genuine biological trend rather than mistaken for one.
Clock selection needs a matching discipline too. The generation and training objective of a clock has to fit what the clinic actually wants to measure. Using a first-generation, age-only clock to check whether a lifestyle intervention is working will likely produce a false negative, because the tool wasn't built to see it, even if the intervention succeeded.
Trajectory modeling itself needs agreed conventions: how many measurements, spaced how far apart, actually constitute a meaningful trend rather than normal week-to-week noise. InCHIANTI's 24-year follow-up proved the concept works, but nobody is running 24-year studies in a clinical setting, so that timeline isn't a template anyone can use in practice.
And frailty-specific validation remains an open gap the field has flagged directly: clocks trained and tested specifically against frailty outcomes, in large, harmonized, longitudinal cohorts, don't yet exist at the scale needed to trust them the way GrimAge is now trusted for mortality. Until these pieces land, epigenetic clocks will keep proving their worth in research cohorts well before they earn a seat at the clinical decision table.
Sources
- Longitudinal changes in epigenetic clocks predict long-term mortality - PubMed
- Longitudinal changes in epigenetic measures over 2 years: methodological implications - PMC
- Longitudinal changes in epigenetic clocks predict long-term mortality | Nature Aging
- Epigenetic clocks: advancing biological age measures towards meaningful clinical use - PMC
- Epigenetic clocks: advancing biological age measures towards meaningful clinical use - eBioMedicine
- Responsiveness of epigenetic aging biomarkers to longevity interventions in humans | Nature Medicine
- fightaging.org
- An unbiased comparison of 14 epigenetic clocks in relation to 174 incident disease outcomes | Nature Communications


