September 4, 2026 · 18 min read · Sugam Budhraja

What Is a Good HRV? Population Studies Say About 40 ms, Fitness Apps Say About 80, and Both Are Right

A 25-year-old man asking what a good HRV is will find 40 ms in a German population cohort, 41 ms in Oura's member data, and 78 ms in WHOOP's. None of those numbers is wrong. They differ because measurement window, device, posture, sleep stage and who got measured all move the number before the person does, which is why the only comparison that controls for all of it is your own baseline.

A 25-year-old man wants to know whether his HRV is good. He finds four answers.

A German population cohort of 1,906 healthy adults says men his age average 39.7 ms [1]. A meta-analysis of 21,438 adults says 42 ms [2]. Oura’s 2024 member data says the average is 41 ms [3]. WHOOP’s published member averages say 78 ms at his age, with the middle half of 20 to 25-year-olds between 55 and 105 [4].

The fitness-app figure is nearly double the population figure. Neither is wrong. They are measuring different things, in different people, at different times of day, with different sensors, and every one of those differences is documented. This post lays them out, because until you know why the numbers disagree, no chart can tell you what a good HRV is.


The four answers, and how each was produced

SourceWhoHow measuredRMSSD for a man in his late 20s
KORA S4 cohort [1]1,906 healthy adults, population-based, Germany5 min, supine, ECG, daytime39.7 ms (mean, ages 25 to 34)
Nunan meta-analysis [2]21,438 healthy adults across studiesShort-term, mostly 5 min, ECG42 ms (mean, all ages)
Lolland-Falster [5]875 adults, ages 15 to 85, Denmark5 min, supine, ECG, daytime27 ms median across all ages
Oura members [3]Members aged 18+, 2024Overnight, finger PPG, 5-min samples averaged40.3 ms mean, 34.2 median (men)
WHOOP members [4]MembersOvernight, wrist PPG, weighted to slow wave sleep78 ms (age 25)

Two things to notice before the explanation.

The three ECG population figures agree with each other once you match the age: about 40 ms for a healthy adult under 35, falling to the high teens by the mid-sixties, which is why Lolland-Falster’s all-ages median lands at 27.

And Oura’s overnight figure lands right on the population figure, while WHOOP’s overnight figure lands nearly double. Both are wrist or finger PPG, both overnight. Something other than “overnight” is doing the work.


Why the numbers disagree

The window is different, and sleep stage changes HRV by more than 100%

Every consumer device reports a number called HRV, and none of them computes it over the same period. In a 536-night validation against ECG, Grosicki and Presby set out exactly what each device was actually averaging [6]:

  • Polar Grit X Pro: the first four hours of sleep only
  • Garmin Fenix 6: the lowest 30-minute average across the entire 24-hour day
  • Oura: five-minute samples averaged across the whole night
  • WHOOP 4.0: “a dynamic average during sleep weighted towards the last phase of slow wave sleep”

Why that matters is a single number from the same paper: normalised high-frequency HRV is 123.7% higher during slow wave sleep than during REM [6]. A device that weights toward slow wave sleep will report a higher figure than one that averages the whole night, from the same heart, with no difference in the sensor at all.

That is most of the gap between Oura and WHOOP in the table above, and it is not a flaw in either. It is two defensible engineering choices producing two non-comparable numbers with the same label.

The device is different, and some are poor at this

Beyond the window, the sensors themselves disagree. Dial and colleagues at Ohio State put five wearables against ECG over 536 nights [7]:

DeviceConcordance with ECGMean absolute errorRating
Oura Gen 40.995.96%Excellent
Oura Gen 30.977.15%Excellent
WHOOP 4.00.948.17%Good
Garmin Fenix 60.8710.52%Poor
Polar Grit X Pro0.8216.32%Poor

Two mainstream devices rated poor for nocturnal HRV. Anyone comparing a Garmin number to an Oura chart is comparing a poor measurement of one window to an excellent measurement of a different one.

Posture alone moves it

Coste and colleagues compared a chest ECG strap against an optical arm sensor in 31 adults, supine and seated. The bias between the two sensors went from -3.21 ms lying down to -8.05 ms sitting up [8]. Same person, same five minutes, same devices, and the act of sitting changed the discrepancy by a factor of 2.5. Recording length, by contrast, barely mattered: two-minute and five-minute recordings gave near-identical results.

The clinical norms above were all measured supine after rest. Almost nobody outside a study takes a spot HRV reading that way.

The people are different

WHOOP’s members are people who bought a device marketed for training and recovery. Elite HRV’s published norms come from 10,308 users of whom 86% were male, and the score is a proprietary 1 to 100 scale rather than milliseconds [9]. Welltory’s chart gives a “normal” rMSSD range of 19 to 107 ms, with 19 to 48 for healthy adults and 35 to 107 for elite athletes, and states outright that “there is no one right HRV number” [10].

KORA and Lolland-Falster sampled the general population. A norm derived from people who track their training is a norm for people who track their training.

Age moves it more than anything, by more than half

This is the factor a single-number chart cannot survive. From the KORA cohort, five-minute supine ECG, mean RMSSD by decade [1]:

AgeRMSSD, womenRMSSD, menSDNN, womenSDNN, men
25 to 3442.9 ms39.7 ms48.7 ms50.0 ms
35 to 4435.432.045.444.6
45 to 5426.323.036.936.8
55 to 6421.419.930.632.8
65 to 7419.119.127.829.6

SDNN is included because Apple Watch reports it rather than RMSSD, so an Apple user reading an RMSSD chart is comparing the wrong metric. A decline of 55% in women and 52% in men on RMSSD across four decades. The steepest single drop is between the mid-thirties and mid-fifties, where both sexes lose over a quarter of their RMSSD in ten years. A 50-year-old with an RMSSD of 25 ms is average. A 25-year-old with the same number is at the low end. The same value has opposite meanings.

On sex, the large cohorts find women slightly higher on RMSSD, with Lolland-Falster reporting men 16% lower [5], though small studies sometimes find the reverse. The difference is real but modest next to age.

Breathing changes the number without changing the physiology

The uncomfortable one. Shaffer and Ginsberg’s overview notes that shifts in respiration rate “can markedly change HRV indices (HF power, RSA, pNN50, RMSSD) without actually affecting vagal tone” [2]. Slower, deeper breathing raises HRV. So a spot reading taken while deliberately relaxing produces a higher number that does not reflect a healthier autonomic system, only a different breathing pattern during the recording. RMSSD is less sensitive to this than some other metrics, which is why it is the short-term standard, but it is not immune.

What does not matter much, for once. Lolland-Falster tested the usual suspects. 53% of participants had caffeine within an hour of measurement, 60% had eaten within three hours, 28% were tested before noon. After correction for multiple testing, none of those showed a significant association with HRV [5]. The coffee is not the problem. The chair, the sensor, the window and the decade are.

Between-person HRV is not meaningless, it is weak

It would be easy to read the above as “HRV cannot be compared between people.” That is not what the evidence says, and the distinction matters.

Hillebrand and colleagues pooled eight cohort studies of 21,988 people without known cardiovascular disease and found that those with the lowest SDNN had a 35% higher risk of a first cardiovascular event than those with the highest (relative risk 1.35, 95% CI 1.10 to 1.67). The dose-response was close to linear: each 1% increase in SDNN was associated with roughly a 1% lower risk [14].

So low HRV means something at population scale. Three qualifications keep that from becoming a chart.

It is SDNN, from long recordings. Shaffer and Ginsberg are explicit that 24-hour SDNN predicts cardiac risk and five-minute SDNN does not [2], and consumer devices report neither: most report overnight RMSSD.

It is a gradient across a population, not a threshold for a person. A relative risk of 1.35 between the extremes of a distribution says nothing about where any individual sits, particularly once the device, window and age effects above have moved their number by a factor of two.

And the effect, while real, is small next to the measurement spread. A 35% risk difference between the top and bottom of the distribution is dwarfed by a 100% difference between two devices reading the same night.

The honest summary. Population HRV carries prognostic information. Population HRV charts do not deliver it, because they are unmatched on the factors that move the number more than the underlying physiology does. The signal is real. The consumer instrument for reading it is not.

So what is a good HRV

The honest population answer exists, and it is narrower than any app chart: about 40 ms RMSSD for a healthy adult under 35, measured for five minutes lying down by ECG, declining to about 20 ms by the mid-sixties. If you want to compare a person to a population, that is the comparison, and it requires matching the device, the window, the posture and the age band, because each of those alone moves the number by 20% to over 100%.

Almost no consumer reading meets those conditions. Which is why the useful answer is a different question.


Why your own baseline is the only comparison that controls for all of it

Every confound above is fixed when the comparison is you-versus-you. Same device, same window, same sleep, same age within a rounding error, same breathing pattern on average. What remains is signal plus day-to-day noise, and the size of that noise is now well measured.

In ordinary adults, it is large. Hannon and colleagues had 41 adults take a supine RMSSD reading every morning for 14 days. The average day-to-day coefficient of variation was 0.37, with individuals ranging from 0.14 to 0.71 [11]. Roughly a third of a person’s own mean, day to day, for no reason a chart would capture.

In elite athletes, it is small. Bellenger and colleagues tracked 11 Olympic water polo players on WHOOP for 16 weeks and found a weekly RMSSD coefficient of variation of 5.4% [12]. Training responses of 10% to 20% therefore stood clearly above the noise for those athletes. For an ordinary adult varying by 37%, a 15% change is inside the noise.

And the baseline needs at least five nights. Grosicki and colleagues analysed roughly two million nocturnal HRV readings from more than 21,000 wearable users and found that at least five of seven nights were required to estimate a week’s day-to-day variation with an intraclass correlation of 0.80 or above [13]. Higher day-to-day variation was associated with greater alcohol consumption, lower physical activity, and shorter, less consistent sleep. The variability itself carries information; a single reading does not.

And sports science has already worked out how to read it. Buchheit’s monitoring review puts the typical measurement error of Ln rMSSD at around 12% expressed as a coefficient of variation, against a smallest worthwhile change of about 3% [16]. A single reading cannot resolve a 3% change inside 12% noise, which is why the review recommends averaging: noise falls by a factor of the square root of the number of readings, and correlations with performance “could only be observed using the average of at least 3 to 4 days,” never from isolated single-day values [16].

Plews and colleagues showed the rolling average in use. Tracking two elite triathletes daily over 77 days, the one who became non-functionally overreached showed a declining 7-day rolling average of Ln rMSSD heading into the event while the control athlete’s stayed flat. The overreached athlete’s day-to-day variability also fell steadily, by 0.65% per week, while the control’s did not move [15]. Two signals: the level of the rolling average, and the variability around it.

Put those together and a rule falls out. One night of HRV tells an ordinary person almost nothing, because their own day-to-day noise is around a third of their mean. Five or more nights establish a baseline. Compare a 7-day rolling average against that baseline, treat a change as meaningful only when it exceeds the person’s own typical error, and watch the day-to-day variability as a second signal, because a fall in variability preceded overreaching in the athletes who were studied. None of that requires knowing what anyone else’s HRV is.
If a spot reading is unavoidable, make it comparable to itself. Lying down, after five minutes of rest, at the same time of day, on the same device, breathing normally rather than deliberately slowly. Two minutes is enough: Coste and colleagues found only marginal differences between two-minute and five-minute recordings [8]. Every one of those conditions is there to hold a confound constant, and the point is not that a supine reading is “correct” but that a supine reading on Tuesday is comparable to a supine reading on Friday.

What this means if you show HRV to users

Show change against the user’s own baseline, not a number against a chart. The chart was built on a different device, window, posture and population, and the reader cannot tell which. Their own baseline controls for all of it at once.

Do not compute a baseline from fewer than five nights, and say so in the interface. A percentile after two nights is a coin flip presented as a finding.

State the window. “Overnight average” and “weighted toward deep sleep” are different quantities with the same name, and the difference exceeds 100% in the underlying signal. A user comparing your number to their partner’s device will otherwise conclude one of you is broken.

If you show a population comparison, match on device, window, age band and sex, or do not show it. Every one of those moves the number by more than most users’ day-to-day change. An unmatched percentile is not a softer claim than an absolute number. It is a less honest one.

Say which metric. Apple reports SDNN, most others report RMSSD, and Shaffer and Ginsberg note that 24-hour SDNN predicts cardiac events while five-minute SDNN does not [2]. The label “HRV” covers at least two quantities that answer different questions.


Where we sit

Sahha reports HRV as a biomarker, uses it in scores, and offers three comparison lenses on it: the user against their own 30-day baseline, the user against a demographic cohort matched on age range and gender, and the user against the global Sahha population. The last two return a percentile. So this post describes a constraint on our own product, and it is worth being precise about which lens it constrains.

The baseline lens is the one the evidence above supports without qualification. Same person, same device, same window. What remains is signal and day-to-day noise, and the rule for separating them is the one in the previous section.

The demographic lens matches age and sex, which removes the two largest physiological confounds. It does not match on device or measurement window, and the central finding of this post is that those move the number by more than age does. A user on a slow-wave-weighted overnight device compared against a cohort that includes five-minute daytime readings is being compared across a gap the percentile cannot see. That lens is useful for “is this normal for someone my age,” and it is honest only when the device and window are held constant or stated.

The global lens matches on nothing, and our own documentation already says to use it for novelty moments and to avoid it for sensitive metrics. On the evidence here, HRV belongs on that avoid list.

Every data log Sahha stores carries its source, recording method and device, which is what makes the first lens trustworthy and the other two improvable. The mechanics are in normalising wearable data across providers, including why an SDNN and an RMSSD must never share a column.

None of that makes a Sahha HRV comparable to a WHOOP HRV. Nothing does. It makes a Sahha HRV comparable to last week’s Sahha HRV for the same person, which is the only comparison the literature supports without a page of caveats.


The short version

“What is a good HRV” has an honest population answer: about 40 ms RMSSD for a healthy adult under 35, measured supine by ECG, falling by more than half by the mid-sixties. Fitness-app charts show higher figures because they measure overnight, weight toward slow wave sleep where high-frequency HRV is over 100% higher, use sensors of varying quality, and sample people who bought a training device.

None of those charts is wrong. None is comparable to any other. And none of them is comparable to a reading you take sitting up after coffee, though the coffee turns out not to matter.

Low HRV does predict cardiovascular risk at population scale, about 35% higher between the extremes, but the charts that claim to read it are unmatched on the factors that move the number most.

Ordinary adults vary by around a third of their own mean from one day to the next. Five nights make a baseline. A 7-day rolling average against that baseline, with change judged against the person’s own noise, is the only HRV signal that controls for every confound in this post at once, and it is the only one worth showing without qualification.

References

  1. Voss A, Schroeder R, Heitmann A, Peters A, Perz S. Short-Term Heart Rate Variability, Influence of Gender and Age in Healthy Subjects. PLOS ONE, 2015. KORA S4 population study, 1,906 healthy subjects, 5-minute supine ECG. Decade values from the age-stratified tables. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0118308
  2. Shaffer F, Ginsberg JP. An Overview of Heart Rate Variability Metrics and Norms. Frontiers in Public Health, 2017. Nunan meta-analysis values from Table 6. https://www.frontiersin.org/journals/public-health/articles/10.3389/fpubh.2017.00258/full
  3. What Is the Average HRV? Oura, member data January to December 2024, members aged 18 and over. Retrieved 4 September 2026. https://ouraring.com/blog/average-hrv/
  4. Good HRV Explained: Average Heart Rate Variability by Age. WHOOP. Figures as published in WHOOP’s article and reported by secondary sources; WHOOP’s site did not serve the page to automated retrieval, so these figures were not read from the page directly. https://www.whoop.com/us/en/thelocker/what-is-a-good-hrv/
  5. Hansen CS, Christensen MMB, Vistisen D, et al. Normative data on measures of cardiovascular autonomic neuropathy and the effect of pretest conditions in a large Danish non-diabetic CVD-free population from the Lolland-Falster Health Study. Clinical Autonomic Research, 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC11937105/
  6. Grosicki GJ, Presby DM. Accurate comparison of wearables requires contextual equivalence. Physiological Reports, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12701519/
  7. Dial MB, Hollander ME, Vatne EA, Emerson AM, Edwards NA, Hagen JA. Validation of nocturnal resting heart rate and heart rate variability in consumer wearables. Physiological Reports, 2025. 13 participants, 536 nights, ECG reference. https://pmc.ncbi.nlm.nih.gov/articles/PMC12367097/
  8. Coste, Millour, Hausswirth. A Comparative Study Between ECG- and PPG-Based Heart Rate Sensors for Heart Rate Variability Measurements: Influence of Body Position, Duration, Sex, and Age. Sensors, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12473955/
  9. Normative HRV Scores by Age and Gender. Elite HRV. 10,308 users with demographic data, 86.1% male. Retrieved 4 September 2026. https://elitehrv.com/normal-heart-rate-variability-age-gender
  10. HRV Chart by Age and Gender. Welltory. Retrieved 4 September 2026. https://welltory.com/hrv-chart-by-age/
  11. Hannon et al. Associations Between Daily Heart Rate Variability and Self-Reported Wellness: A 14-Day Observational Study in Healthy Adults. Sensors, 2025. 41 participants, Polar H10, supine RMSSD on waking. https://pmc.ncbi.nlm.nih.gov/articles/PMC12300306/
  12. Bellenger et al. Evaluating the Typical Day-to-Day Variability of WHOOP-Derived Heart Rate Variability in Olympic Water Polo Athletes. Sensors, 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC9505647/
  13. Grosicki GJ, Carter JR, Laursen PB, Plews DJ, Altini M, et al. Heart rate variability coefficient of variation during sleep as a digital biomarker that reflects behavior and varies by age and sex. American Journal of Physiology, Heart and Circulatory Physiology, 2026. https://journals.physiology.org/doi/full/10.1152/ajpheart.00738.2025
  14. Hillebrand S, Gast KB, de Mutsert R, Swenne CA, Jukema JW, Middeldorp S, Rosendaal FR, Dekkers OM. Heart rate variability and first cardiovascular event in populations without known cardiovascular disease: meta-analysis and dose-response meta-regression. Europace, 2013. Eight studies, 21,988 participants. https://academic.oup.com/europace/article/15/5/742/673395
  15. Plews DJ, Laursen PB, Kilding AE, Buchheit M. Heart rate variability in elite triathletes, is variation in variability the key to effective training? A case comparison. European Journal of Applied Physiology, 2012. https://pubmed.ncbi.nlm.nih.gov/22367011/
  16. Buchheit M. Monitoring training status with HR measures: do all roads lead to Rome? Frontiers in Physiology, 2014. https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2014.00073/full

Related