MoreYears reader guide

How to read longevity evidence

A result can be exciting without being ready to guide a decision. This ladder helps separate biological possibility from evidence about what happens in people.

The evidence ladder

Five rungs, five different questions

Research does not always move neatly upward, and a study’s design does not guarantee its quality. The rungs describe what kind of conclusion a method can support—not whether a particular result is automatically trustworthy or important.

  1. Cells and mechanisms

    Biological plausibility

    What it can tell us

    Whether a process can change in cells or tissues, and which biological pathways may be involved.

    Central limitation

    A controlled cellular effect does not show that the same approach will be effective or safe in a whole organism.

  2. Animal research

    Whole-organism testing

    What it can tell us

    How an intervention behaves across organs and systems under controlled conditions that may be impossible in people.

    Central limitation

    Species biology, doses, environments, and disease models can differ substantially from human aging and clinical use.

  3. Human observational studies

    Patterns in people

    What it can tell us

    Whether an exposure, behavior, biomarker, or outcome tends to occur alongside another in real human populations.

    Central limitation

    Confounding and reverse causation can create an association even when one factor does not cause the other.

  4. Randomized human trials

    Causal effects under trial conditions

    What it can tell us

    Random assignment can estimate whether an intervention caused a difference for the people, outcomes, and timeframe studied.

    Central limitation

    A trial may be too small or short, use a surrogate endpoint, miss rare harms, or not generalize beyond its participants.

  5. Replicated and synthesized human evidence

    Consistency across studies

    What it can tell us

    Whether findings hold up across independent studies, settings, populations, or a careful synthesis of the available evidence.

    Central limitation

    A synthesis inherits weak studies, inconsistent methods, publication bias, and limits on personal applicability.

Across every rung

Study design is only the beginning

Two studies on the same rung can deserve very different weight. These questions help reveal whether a result is durable, meaningful, and relevant.

Replication and convergence

Has an independent team found a similar result? Confidence grows when different methods point in the same direction, not merely when one study is repeated in the same setting.

Effect size and uncertainty

Is the difference large enough to matter, and how wide is the plausible range around it? Statistical detection alone does not establish practical importance.

Population relevance

Age, health, sex, ancestry, baseline risk, and selection criteria shape whom a result describes. Evidence in one group may not transfer cleanly to another.

Endpoints

A biomarker, symptom, physical function, disease event, healthspan, and lifespan answer different questions. Movement in one should not be silently translated into another.

Duration

Short studies may detect an early signal while missing whether it lasts, whether benefits accumulate, or whether delayed harms emerge.

Conflicts and funding

Funding or commercial involvement does not automatically negate a result, but readers should know who designed, analyzed, and reported the work.

Peer review

Peer review adds scrutiny but is not a guarantee that methods are sound, reporting is complete, or later studies will agree.

Preprints

Preprints can share results quickly, but they have not completed journal peer review and may change materially before publication.