September 14, 2026 · 11 min read · Sugam Budhraja

What Is a Good Sleep Score? Apple's 82 Is High, Garmin's 82 Is Good, and WHOOP Is Scoring Something Else

Apple's sleep score is 50% duration, 30% bedtime consistency over 13 nights, 20% interruptions. Garmin folds in overnight stress from HRV. WHOOP Sleep Performance is four factors, not hours divided by need. The same 82 is a different sentence on each device, and the 0–100 grade has not been validated end-to-end against a lab or clinical reference.

A user asks whether last night was good. Four products answer with a number that looks interchangeable.

On Apple Watch, 82 is High [1]. On Garmin, 82 is Good — Excellent starts at 90, and Garmin’s own users averaged a Fair 71 in 2024 [2]. Oura also speaks 0–100, but that 82 is built from stages, efficiency and timing that Apple does not even put in the formula [3]. WHOOP may show 82%, which is not a quality grade: Sleep Performance is four factors, and hours-versus-need is only one of them [4].

Those are statements about scales, not about one night that produced four 82s. The 0–100 grade has not been tested end-to-end against a lab or clinical reference. Sleep/wake and stages have. A good sleep score is a band inside one product. This post is the map of those bands, the formulas that are public, one reconstructed night where the inputs can be traced, and the reason chasing 90 is a product mistake.

This is not Sahha’s Sleep Score walkthrough. The seven factors, the phone-only path, and how to use the breakdown in a product live in the Sleep Score guide. This piece is the public question the guide is not trying to rank for: what “good” means when every vendor prints 0–100 on a different construct.

What does each vendor actually score?

Short answer: Apple scores duration, 13-night consistency and interruptions, with published weights. Garmin adds overnight stress. Oura adds stages and timing. WHOOP scores four named factors, not hours divided by need. Only Apple publishes the weights. That is not the same as publishing how those 50, 30 and 20 points are awarded inside each block.

ScoreScaleWhat it is built fromWhat “good” means on that scale
Apple Sleep Score [1]0–100Duration 50 (how long you slept), bedtime consistency 30 (last 13 nights), interruptions 20Very Low 0–40, Low 41–60, OK 61–80, High 81–95, Very High 96+
Garmin Sleep Score [2]0–100Duration plus quality: stages, awake time, restlessness, and average stress during sleep from HRV (Firstbeat)Excellent 90–100, Good 80–89, Fair 60–79, Poor below 60. 2024 user average: 71 (Fair)
Oura Sleep Score [3]0–100Seven contributors: total sleep, efficiency, restfulness, REM, deep, latency, timingLabels on a 0–100 construct (Optimal toward the top). Not Apple’s 0–100.
WHOOP Sleep Performance [4]0–100%Sufficiency (hours versus estimated need), consistency versus the prior four days, efficiency, sleep stress100% is the top of that four-factor mix, not “perfect architecture”

Apple is the honest outlier on transparency. The Sleep app still shows stages and heart-rate context; they do not enter the 0–100 [1]. A long, regular, uninterrupted night can score Very High with almost no deep sleep on the stage chart. That is not a bug. It is the formula.

Garmin is the honest outlier on physiology. Overnight stress from HRV means a 7.5-hour night with a sympathetic load — alcohol, illness, late training — scores worse than the same hours on Apple, which cannot see that load in the composite [2].

WHOOP is the honest outlier on the question. Sufficiency is hours versus a personal need; Sleep Performance is that plus consistency, efficiency and overnight stress [4]. Two people who slept 7 hours can land far apart because the model thought one of them needed 8.5, or because one of them went to bed two hours late. Reducing the percentage to hours ÷ need is the error the reconstructed night is for.


Why the same 82 means four different things

The construct is different

Put the same sleeper on an Apple Watch and an Oura Ring and you already know the stage minutes will disagree — Robbins and colleagues, comparing three wearables against polysomnography in 2024, found Apple Watch overestimated light sleep by 45 minutes and underestimated deep by 43, while Oura did not differ significantly from the lab [5]. We unpacked that in Apple Watch vs Oura. A score that includes REM and deep sleep inherits that disagreement. Apple’s score, which ignores stages, does not. That is why the two 82s can move in opposite directions on the same night: one is reacting to a stage estimate the other does not use.

The band is different

Apple’s High starts at 81 [1]. Garmin’s Good starts at 80, but Excellent — the word users hear as “good” — starts at 90, and Garmin’s published 2024 average was a Fair 71 [2]. In late 2022, 5% of Garmin sleep users averaged in the Excellent range [2]. That is a statement about people over three months, not a claim that 95% of nights miss 90. Designing a product that congratulates 90+ is designing for Garmin’s tail, then applying the same cheer to Apple, where 82 is already High.

The date is different

Even duration, the one input every score uses, is not the same object. Fitbit, Oura, Garmin, WHOOP, HealthKit and Health Connect do not agree on which calendar day last night belongs to. A composite that includes “last night” will jump when the user switches source, for a reason that has nothing to do with sleep.

The composite is untested

Validation studies of wearables almost always stop at sleep/wake and stages. They do not then test whether Oura’s 85 predicts next-day performance better than Apple’s 85. Human scorers looking at the same polysomnography recording already disagree on stages; we covered the ceiling in where accuracy claims actually come from. A 0–100 that folds those stages in, then refuses to publish weights, is a headline. It is not a biomarker.

One night, four subsets

The composites are unpublished except for Apple’s weights, so nobody can emit Oura’s 0–100, Garmin’s Firstbeat number, or WHOOP Sleep Performance from the physiology alone. What we can do is hold the night still and say which listed inputs move.

Take a constructed night, labeled as such:

  • 7 hours asleep, 7 hours 40 minutes in bed (~91% efficiency)
  • Sleep onset at 01:00, against a 13-night median of 23:00
  • Apple Health sleep goal of 8 hours
  • WHOOP sleep need of 8 hours 30 minutes (sufficiency = 7 ÷ 8.5 ≈ 82% — that fraction is not Sleep Performance)
  • One 25-minute awakening
  • A hard evening session, so overnight HRV/stress is off-baseline
ScoreInputs that move on this nightInputs that do not enterWhat we can know
Apple Sleep ScoreDuration is 7 hours, short of an 8-hour Health sleep goal the app shows next to time asleep, so the 50-point duration block has room to move. Consistency versus the last 13 nights — onset two hours off the median is the 30-point hit. One 25-minute wake moves interruptions [1]Stages, HRV, efficiency as a named factorDirection of the three published weights. Not an 82
Garmin Sleep ScoreDuration short of typical; overnight stress from HRV after the evening session is the differentiator versus Apple [2]Apple’s 13-night consistency as a 30-point blockStress-during-sleep can pull Firstbeat when Apple cannot see it. Not a 0–100
Oura Sleep ScoreTotal sleep, timing, restfulness from the awakening; efficiency ~91% is usually fine. REM and deep are unknown to us and, on Apple, unused [3]Apple’s 50/30/20 splitTiming and total sleep move. Stages Apple ignores. Not a 0–100
WHOOP Sleep PerformanceSufficiency ~82% of need. Consistency versus the prior four days (the late night dings it). Efficiency ~91%. Sleep stress from the evening session [4]Apple’s 13-night windowSufficiency is 82%. Sleep Performance is four factors. It is not 82%

The useful output is not a fake 82 / 82 / 82%. It is that the same night is a different subset of the same sleep. A product that charts four composites as confirmation is charting four questions. A product that shows duration, timing and continuity is showing the night. That split is also why this page and the readiness companion have to stay separate: overnight stress belongs in readiness more than it belongs in “how well did I sleep.”


What the clinical literature will actually support

If you strip the branded 0–100 away, the evidence that remains is not a score. It is a small set of dimensions.

Duration. Adults sleeping 7 to 9 hours sit at the low point of a U-shaped curve for mortality and cardiometabolic risk; under 6 hours and over 9 both rise [6][7]. That is the 50 points in Apple’s formula, and it is the one input a phone can estimate without a Watch.

Regularity. In the Multi-Ethnic Study of Atherosclerosis, Huang and colleagues followed 1,992 adults with seven days of wrist actigraphy. Compared with people whose sleep-onset time had a seven-day standard deviation of 30 minutes or less, those whose onset SD exceeded 90 minutes had a hazard ratio of 2.11 for incident cardiovascular events. Duration irregularity was similar: SD above 120 minutes versus 60 or less had a hazard ratio of 2.14, independent of average duration [8]. That is a week of scattered timing, not one bedtime that wandered 90 minutes. Apple still put 30 points on 13-night consistency. Oura scores timing as a contributor. A product that only praises 8 hours at 2am is ignoring the half of the literature that is not hours.

Continuity. Fragmented sleep is not just “a worse night.” In a pooled analysis of 8,001 older adults, nocturnal arousal burden on polysomnography was associated with long-term cardiovascular and all-cause mortality, more clearly in women than in men [9]. Apple’s remaining 20 points are interruptions. Garmin’s restlessness and awake time are the same idea with HRV mixed in.

Those three are the honest core. Stages are directional. HRV-during-sleep is a different construct again, closer to readiness than to “how well did I sleep.” Mixing them into one 0–100 is a product choice. It is not a finding.


What a product should put on screen

Do not compare scores across devices. An 82 on Apple and an 82 on Garmin are not a confirmation. They are a coincidence. If the user wears two sensors, pick a primary source or drop the composite and show duration, timing and continuity, which you can reconcile.

Do not set 90 as a goal. Garmin’s 2024 average was 71 [2]. Excellent starts at 90. In late 2022, 5% of Garmin sleep users averaged in that band — a fact about people over three months, not a nightly miss-rate [2]. Apple’s Very High starts at 96. A goal parked in a vendor’s tail is how you buy orthosomnia in the name of engagement.

Show the factors that moved. Apple already does this, because the formula is three numbers. Oura does it with seven contributors. A single grade without a because is how users decide the product is lying when the number and the feeling disagree.

Prefer constructs that survive a missing wearable. Most of the people you want to score do not wear a Watch to bed. Duration, regularity and debt can be estimated from the phone; stages cannot, and overnight HRV cannot. That split is the whole point of the Sleep Score guide. The blog’s job is to stop you importing Apple’s 82 into a cross-platform chart and calling it science.

A good sleep score is the band your vendor printed next to a construct they defined. The useful object underneath is still last night’s length, timing and continuity. Show those, and the 0–100 is optional. Show only the 0–100, and you have shipped someone else’s recipe with your logo on the grade.

References

  1. Apple. View your sleep score on Apple Watch. Apple Support. https://support.apple.com/guide/watch/view-your-sleep-score-apded441a669/watchos
  2. Garmin. How do Garmin watches track your sleep, calculate sleep score. Garmin Blog. https://www.garmin.com/en-SG/blog/general/how-do-garmin-watches-track-your-sleep/
  3. Oura. Sleep Score. https://ouraring.com/blog/sleep-score/
  4. WHOOP. How well WHOOP measures sleep. The Locker. https://www.whoop.com/us/en/thelocker/how-well-whoop-measures-sleep/
  5. Robbins, R., Weaver, M. D., Sullivan, J. P., et al. (2024). Accuracy of three commercial wearable devices for sleep tracking in healthy adults. Sensors, 24(20), 6532. https://doi.org/10.3390/s24206532
  6. Cappuccio, F. P., D’Elia, L., Strazzullo, P., & Miller, M. A. (2010). Sleep duration and all-cause mortality: a systematic review and meta-analysis of prospective studies. Sleep, 33(5), 585–592. https://doi.org/10.1093/sleep/33.5.585
  7. Itani, O., Jike, M., Watanabe, N., & Kaneita, Y. (2017). Short sleep duration and health outcomes: a systematic review, meta-analysis, and meta-regression. Sleep Medicine, 32, 246–256. https://doi.org/10.1016/j.sleep.2016.08.006
  8. Huang, T., Mariani, S., & Redline, S. (2020). Sleep irregularity and risk of cardiovascular events: the Multi-Ethnic Study of Atherosclerosis. Journal of the American College of Cardiology, 75(9), 991–999. https://doi.org/10.1016/j.jacc.2019.12.054
  9. Shahrbabaki, S. S., Linz, D., Hartmann, S., Redline, S., & Baumert, M. (2021). Sleep arousal burden is associated with long-term all-cause and cardiovascular mortality in 8001 community-dwelling older men and women. European Heart Journal, 42(21), 2088–2099. https://doi.org/10.1093/eurheartj/ehab151

Related