The Evolution Index · Methodology Home
Matt Roberts Evolution · Protocol v1.1 · Provisional

The method behind the number

The Evolution Index expresses a person's functional capacity as a single number from 0 to 100 — normed to their age and sex, alongside a band and a Capacity Age. It exists to make decline visible early enough to reverse, and to direct effort at the one input that will move the most. This is how it works, and why we trust it.

Get your Provisional Index Book the full assessment

The Provisional Index is free. You enter the measurements yourself and you create your account on that page. It needs numbers a watch or a gym can give you — a VO₂ max estimate, grip, body composition — and it tells you how much each estimate widens the error. The certified assessment is taken in person at 32 Grosvenor Square on calibrated equipment. If you have none of that to hand, the Field Test needs only a staircase and a chair.

Four principles

Measurement is not training

The Index exists to direct work, not replace it. Every report ends with one intervention, never a list.

Domains do not compensate

Cardiorespiratory fitness, strength and movement are independent risk pathways. Excellence in one does not offset failure in another — and the scoring enforces it.

Nothing beyond the evidence

Where a measure is included for engagement rather than hard outcome evidence, we say so — here, and in the report. Where a measurement device cannot carry the claim, the value is recorded as context and does not score.

Error is published

Every metric carries a stated minimum detectable change. Below it, we report no change. Almost nobody in this market publishes their error. We do.

What we measure, and why

Three domains, weighted by the strength of the evidence linking each to all-cause mortality and loss of independence. The weighting is deliberate: cardiorespiratory capacity carries the most, because the evidence behind it is the strongest.

40%
Cardiorespiratory

How well the body takes in, transports and uses oxygen — the single most powerful predictor of longevity in the battery.

  • VO₂ max — the dominant input. Maximal oxygen uptake, the gold-standard measure of aerobic capacity. Evidence-anchored
  • Heart-rate recovery — how fast the heart settles after maximal effort, a window on autonomic health. Evidence-anchored
35%
Strength & tissue

The muscle and connective tissue that carry you through life and protect you when things go wrong. Four measures in v1.1: grip and lean tissue carry equal and largest weight, then lower-body power, then the hang.

  • Grip strength — a proxy for whole-body strength and a robust mortality signal. Evidence-anchored
  • Fat-free mass index (FFMI) — whole-body lean tissue scaled to height; the tissue sarcopenia takes. Scored only when the acquisition method can carry the claim. Evidence-anchored
  • Lower-body power — jump, five-repetition sit-to-stand, or the single-leg sit-to-stand from graded seat heights; power fades earlier and faster than strength, so it warns first. Function-validated
  • Dead hang — grip endurance under load, expressed as work. Engagement-validated
25%
Movement capacity

The ability to get down to the world and back up from it — mobility, balance and control.

  • Sit-rise test — getting to the floor and up with as little support as possible. Evidence-anchored
  • Single-leg stand — static balance, a documented predictor of survival. Evidence-anchored
  • Deep squat hold — end-range strength and mobility, held over time. Function-validated
Evidence-anchored — direct outcome evidence Function-validated — measures the capacity directly Engagement-validated — tracks anchored measures; included honestly as such

How each test is run

Standardisation is everything: the same conditions, the same scripted words, the same equipment at every retest — because a change in the test can masquerade as a change in the person. Each test below lists what it captures, how it is administered, and its published error — the minimum detectable change (MDC) below which we report no change at all.

VO₂ max

Ramp or Bruce protocol to volitional exhaustion on treadmill or cycle, with the modality fixed for the client's history. A true maximum is confirmed against objective criteria; where it can't be, it is recorded as a peak and flagged.

Published error · MDC 2.0 ml/kg/min

Heart-rate recovery (60s)

A complete stop. At the instant the maximal test ends the client comes to a full halt and stands motionless — no walking cooldown, no continued pedalling, no sitting, no hands on knees, no leaning on the rail. The timer starts at the moment of cessation. The score is peak heart rate minus heart rate at exactly 60 seconds. Posture and cessation dominate this measure, so both are scripted and held identical at every retest; any deviation is recorded and the value flagged rather than quietly compared.

Published error · MDC 6 bpm

Grip strength

Seated, elbow at 90°, a calibrated hydraulic dynamometer. Three maximal efforts per hand with rest between; the best dominant-hand value is used. The same device is used at every retest — brands are not interchangeable.

Published error · MDC 4 kg

Fat-free mass index (FFMI)

Whole-body fat-free mass in kilograms divided by height in metres squared, measured fasted and normally hydrated on the same device each time. FFMI replaces appendicular lean mass in v1.1: it is available from every reference-grade body-composition method rather than DEXA alone, it is more robust to limb-segmentation differences between scanners, and it is what the devices in the field actually report.

The acquisition method decides how the value is used, because a number is only as good as the instrument that produced it:

  • DEXA or BodPod / air displacementCertifiable scores, and the record may be certified.
  • Multi-frequency BIAProvisional scores, but the record cannot be certified and is excluded from norms.
  • Single-frequency BIA, skinfolds, estimatesContext only recorded and shown, never scored. The weight redistributes across the rest of the domain.
Published error · MDC 0.5 kg/m² (DEXA / BodPod)

Lower-body power

One of three, fixed for the client's history. Countermovement jump — hands on hips, best of three, height in centimetres — for the private cohort. The timed five-repetition sit-to-stand is the validated alternative for corporate and older clients. New in v1.1, the single-leg sit-to-stand is the most discriminating of the three: the client rises from a graded seat, unaided, on the weaker leg, and the score is the lowest seat height they can clear. Seats are stepped in fixed increments and the same stack is used at every retest.

The single-leg version replaces the farmer's carry in the battery. It isolates one limb, so it exposes the asymmetry a bilateral test conceals; it needs a box and nothing else; and unlike a loaded carry it does not confound leg capacity with grip, which the battery already measures twice. Power is a more sensitive early-warning signal than maximal strength.

Published error · MDC 2 cm (jump) · 1.2 s (5×sit-to-stand) · 4 cm (single-leg)

Dead hang

A passive hang from a fixed-diameter bar, feet clear, to release. Scored as work — hang time multiplied by bodyweight — so that heavier clients are not penalised for carrying more tissue. Single attempt, capped for safety.

Published error · MDC 15%

Sit-rise test

Barefoot, on a clear floor: sit down and stand back up using as little support as possible. Points are deducted for each hand, knee or loss of balance used. Deduction rules are posted and one assessor owns a client's scoring across their history.

Published error · MDC 1.0 point

Single-leg stand

Eyes open, hands on hips, timed to the first loss of position, best of two. The published mortality evidence uses a 10-second pass/fail; we use a timed scale, with 10 seconds sitting near the base of the range.

Published error · MDC 5 s

Deep squat hold

Barefoot, hips below the knee crease, heels flat, torso self-supported, held for time. Continuous and directly trainable — it represents the very capacity the domain claims to measure.

Published error · MDC 10 s

How the score is built

Every raw measurement — a VO₂ figure, a grip in kilograms, a balance time — converts to a 0–100 metric score through the same anchor system, so that a score means the same thing on every test. A 55 is a 55 whether it came from a treadmill or a squat hold.

0Floor of measurable capacity
40Entry to Slipping
55The median for your age & sex — entry to Holding
70Entry to Resilient
85Entry to Extending
100Ceiling

Scores between anchors are interpolated smoothly; for tests where lower is better — the timed sit-to-stand, the seat height in the single-leg rise — the scale simply runs the other way. Thresholds are set at a reference age and adjusted continuously for the individual, so nobody gains or loses points overnight on a birthday.

The bands

85–100 · ExtendingCapacity ahead of chronology and still being extended.
70–84 · ResilientWell protected. Absorbs illness, injury and interrupted training without losing ground.
55–69 · HoldingAbove median, no margin. Decline resumes when the inputs stop.
40–54 · SlippingMeasurably losing capacity. Fully reversible.
0–39 · ReversingActively reversing into decline. Independence is at stake.

The top band is named Extending in v1.1. The previous name, Compounding, borrowed a financial metaphor that implied returns accrue on their own. They do not: capacity at this level is held and pushed further only by continued work, and the name now says so.

The non-negotiable floor

If any single domain scores below 40, the overall Index is capped at 59 — the top of Holding — no matter how strong the arithmetic elsewhere. A superb VO₂ max cannot buy back sarcopenic muscle. This is the methodological expression of the central claim: these are independent pathways, and they do not trade against one another. We never apply the cap silently — the report states the uncapped score, the cap, and the domain responsible.

Capacity Age

Capacity Age is the chronological age at which your current results would be typical — the age at which they'd sit right at the median. It's the translation device: one number for precision (the Index), one for meaning (the age). It is a functional statement and nothing more — not a biological age, metabolic age or heart age. That discipline is deliberate, and it is what keeps the measure honest.

The age model, and the 18–29 group

Thresholds are published at a reference age of 55 and scaled continuously from there, so the norms move with the person rather than stepping on a birthday. Lean tissue is scaled on a gentler curve than the performance tests: whole-body fat-free mass carries bone and organ tissue, which does not fall away at the rate limb muscle does, and a multiplicative curve borrowed from the strength tests would make the FFMI standard implausibly easy in the eighth decade.

Version 1.1 adds an 18–29 age group, and with it a correction to the model. Below 30 the norms plateau: an 18-year-old and a 30-year-old are scored against the same expectation. Extending the linear curve below 30 implies that capacity keeps rising through the early twenties at the same rate it falls after 55, which the population data does not support — peak capacity is reached in the late twenties and holds. Without the plateau a fit 22-year-old would be scored against a standard nobody has measured, and would be marked down for it.

Capacity Age is floored at 30 for the same reason. Below the plateau the norms are flat, so a Capacity Age of 24 would be an artefact of the arithmetic rather than a finding. Results at or beyond the plateau are reported as <30, and the Index continues to separate performance above that point.

Enhancement inputs: resting HR and HRV

Resting heart rate and heart-rate variability are recorded where a client has them, from a wearable or a morning reading. They are enhancement inputs: they never move the Index. Not one point, in either direction.

What they do is narrow the confidence interval reported around the Index. The composite carries a published interval of ±3 points. A single test session is a snapshot, and a snapshot of a person who slept badly, travelled, or arrived under-fuelled is a noisier estimate of their capacity than the same session backed by months of autonomic data. Continuous resting data tells us how representative the day was — so it buys precision, not score.

The narrowing is weighted by how much history a reading represents:

  • 90-day average — the full narrowing credit.
  • 30-day average — most of it.
  • 7-day average — a modest amount.
  • Single reading — nearly none; one morning's number is a data point, not a baseline.

With both inputs supplied at 90 days the interval narrows to ±2.0 points, and no further — the remaining uncertainty belongs to the tests themselves, and no amount of wearable data removes it.

What narrowing does not do

A narrower interval does not lower the threshold at which we report a change. The composite minimum detectable change stays at 4 Index points for every client, whatever their wearable history, until the test–retest study on the validation roadmap replaces the estimate with a measured figure. Tightening the reporting rule on the strength of data that never entered the score would be exactly the kind of quiet inflation this protocol exists to avoid.

How the norms were built

The scale is anchored on established clinical and population science, then extended by Evolution's own assessment experience where the literature runs out:

  • Sarcopenia cut-points for grip and lean mass (EWGSOP2).
  • Fat-free mass index reference distributions from population body-composition surveys (Schutz and colleagues; Kyle and colleagues).
  • Cardiorespiratory norms (Cooper Institute, ACSM).
  • The sit-rise test (Brito and colleagues) and single-leg balance (Araujo and colleagues).
  • Evolution assessment experience for the newer tests — the single-leg sit-to-stand, the dead hang and the deep squat hold — which carry the widest uncertainty and are labelled accordingly.
What we don't publish — and why

The exact cut-points that convert a raw result into a score, the age-scaling that personalises them, and the precise weightings within each domain are Evolution's proprietary calibration. The method above is open — anyone can inspect and reproduce the logic, and that openness is what makes the Index citable. What is ours is the calibration those norms are tuned to, sharpened against the Evolution client dataset. You are not being asked to trust a black box; you are being shown the whole machine, minus the settings we earned.

Honesty about error

Every metric has a stated minimum detectable change, and the composite Index carries an error of ±3 points — narrowing to no less than ±2.0 where enhancement inputs are supplied, as described above. A change of four points or more between assessments is real and reportable; anything smaller is stated as unchanged, with the reason given. Retest at twelve weeks, not sooner — the delta is the product, and a delta smaller than the error is not a delta.

Where this is going

The Index is publishable now as a provisional instrument. It becomes progressively more defensible — and more genuinely proprietary — through a deliberate validation roadmap:

  1. Retrospective fit — run the existing client dataset through the model and confirm that a score of 55 really is the median, adjusting where the data disagrees.
  2. Fitted age curves — replace the interim linear age model with curves fitted to the Evolution dataset, and test the below-30 plateau against the young cohort as it accumulates. This is the point at which the norms become ours rather than borrowed.
  3. Empirical error — test–retest on a cohort to replace estimated error bars with measured ones, including whether autonomic history genuinely predicts session-to-session variance at the size v1.1 assumes.
  4. Prospective tracking — follow the Index against outcomes that matter: falls, fractures, admissions, sick days. A multi-year asset, and the foundation of anything peer-reviewed.

A scoring system nobody can inspect gets ignored. One that can be inspected — and that states plainly where it is still provisional — gets trusted, and eventually cited. That is the whole idea.