Validating a health score means comparing its output against an external reference standard, in a study, using ground-truth labels the score did not have access to. Everything else that gets called validation is something weaker.
That distinction matters because a vendor telling you their scores are validated could mean any of four different things, and the gap between the weakest and the strongest version is the difference between a marketing decision and a research programme.
This is not pedantry. In August 2026, a proposed class action was filed against Oura over the gap between a number and the task that number describes. The allegations are unproven. But the underlying confusion is everywhere in this category, and it is worth being able to read a claim precisely, including ours.
Four claims, one word
| Level | The claim | What it proves | What it does not |
|---|---|---|---|
| 1. Open | The algorithm is public | Exactly what it computes | Whether the output corresponds to anything |
| 2. Expert-built | A credentialed scientist designed it | The construct is not arbitrary | That the output was ever measured against reality |
| 3. Evidence-based | Each component is grounded in published research | Every factor and threshold is defensible | That the composite was tested end to end |
| 4. Validated | Output tested against a reference standard, in a study | The output corresponds to something external | That it generalises beyond that population and task |
Most of this industry, including us on most scores, lives at level 3. Very little of it reaches level 4 for composite wellness scores. The problem is not that level 3 is bad. It is a genuinely strong position. The problem is that all four levels get described with the same vocabulary.
What open source proves
Publishing an algorithm is a real contribution and the recent move toward it in this category is good for buyers. Be clear about what it buys you.
It proves what the algorithm computes. You can read the thresholds, check the weighting, see whether the documentation matches the code, and confirm that nothing surprising happens to your users’ data on the way to a number. That is genuine, and it is more than most vendors offer.
It cannot prove the output means anything, because the thing that would prove that is not in the repository. Correctness here is not a property of code, it is a property of the relationship between the code’s output and the world, and you cannot read that relationship out of a source file. An open algorithm can be perfectly auditable and still be measuring the wrong construct. A closed one can be rigorously validated.
Openness and accuracy are close to orthogonal. Treating them as the same axis is the single most common error in reading these claims, and it is an easy one to make because transparency feels like honesty and honesty feels like correctness.
What “validated by a scientist” usually means
A credentialed researcher designing or reviewing an algorithm is worth having. It means thresholds came from the literature rather than from product intuition, and it is a meaningful step above assembling a score from guesses.
It is expert review, not measurement validation. The distinction is whether anyone compared the output to something external, with participants, a protocol and an analysis plan. A neuroscientist choosing a defensible HRV window is doing level 2 and possibly level 3 work. Neither becomes level 4 by virtue of the credential.
This matters because the formulation “designed and validated by [expert]” is common, reads as level 4, and usually is not. That is rarely deliberate. Validated is simply a word that has drifted.
Even real validation has a ceiling
The uncomfortable part, and the reason level 4 is harder than it sounds: the reference standard is not perfect either.
Polysomnography is the gold standard for sleep staging. A multicentre validation of eleven consumer trackers against PSG across 349,114 epochs found Cohen’s kappa for stage agreement ranging from 0.06 to 0.56, with no device reaching substantial agreement. That looks damning until you see the control: trained human scorers reading the same recording reach kappa around 0.57 on deep sleep and 0.24 on N1.
So the ceiling on agreement is not set by the wristband. It is set by the fact that two qualified humans reading identical data disagree substantially about what stage someone was in. A device cannot exceed the reproducibility of the thing it is being measured against.
Which is why a bare percentage is uninterpretable. Any validation number needs four things attached: the task (sleep versus wake, or four-stage classification), the metric (raw agreement, kappa, sensitivity and macro F1 give different numbers from identical data), the population (healthy young adults or a clinical cohort), and the reference (one scorer, a consensus, or an automated system).
The reporting standard almost nobody cites
For sleep specifically, a voluntary standard already exists. ANSI/CTA/NSF-2052, developed with the National Sleep Foundation, covers terminology for wearable sleep monitors and the methods for measuring their performance. It sets two separate compliance categories, one for sleep and wake determination and one for sleep stage determination, and asks that performance be reported as mean absolute difference together with a measure of its variability rather than as a single point estimate.
It is cited remarkably rarely. For composite wellness, readiness or resilience scores outside sleep, no comparable reporting standard exists at all, which is part of why the vocabulary drifted in the first place.
Where Sahha actually sits
Applying the ladder to ourselves, because a post like this is worthless otherwise.
The behavioural and mental health work reaches level 4. That rests on a longitudinal study rather than an argument: participants recruited through Prolific in waves from February 2022 with collection through mid-2025, passively monitored for five to twelve or more weeks, completing validated psychometric instruments weekly to provide ground-truth labels, under University of Otago Human Ethics Committee approval, reference 21/074. The study protocol is published, including its limitations. If you are comparing vendors, that reference number is the thing to ask every one of them for.
Score construction is level 3, which is where the strong version of this argument actually lives. Every factor and every goal in a Sahha score is grounded in published research, and our score accuracy guide sets this out factor by factor: seven to nine hours associated with lowest mortality risk, irregular sleep timing carrying risk independent of duration, circadian misalignment tied to metabolic and cardiovascular outcomes. A composite built from individually evidenced factors, each one visible and separately scored, is a defensible construct and a more useful one than an opaque number with a percentage attached.
What it is not is a claim that the composite sleep score has been validated against polysomnography, and we do not make that claim. Nor, as far as we can find, does anyone else in this category make it with a citation attached. Ground truth for depression, anxiety and stress does not transfer into a validated sleep-stage claim, and a company that ran one study does not get to describe its whole product as validated. We would rather draw that line ourselves than have a prospect find it.
The other thing worth saying: our scores return null rather than guessing when an input is missing, and phone-only factors are labelled separately from wearable-dependent ones. Returning nothing is a validity decision, not a gap in coverage.
Six questions for any vendor, including us
- What reference standard was the output compared against, if any? If the answer is none, that is fine, but it should be said rather than implied.
- Who were the participants, and how many? Thirty-five healthy adults for one night is a different claim from a year of longitudinal data.
- What metric is the number you quote, and for which task? A figure for sleep-versus-wake is not a figure for staging.
- What does the score return when an input is missing? A score that always produces a number is telling you less than one that sometimes refuses.
- Was there ethics approval and a pre-specified analysis? This separates a study from a retrospective look at data you already had.
- Can I see the factor breakdown behind a single score? If a number cannot be decomposed, no claim about it can be checked.
A vendor answering all six is being straight with you regardless of how flattering the answers are. A vendor answering none while using the word validated is worth pressing, and the pressing is cheap: every one of these is a question, not a procurement exercise. If you are running a wider evaluation, our comparison pages carry the dated, sourced version of what each platform publishes, and the science guides set out the factor-level evidence behind each Sahha score.
The honest summary of this category in 2026 is that almost nobody is at level 4 for composite scores, several are at level 3, a growing number are at level 1, and the marketing language does not reliably distinguish them. Knowing which rung a claim is standing on is most of the work.
References
- Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study. JMIR mHealth and uHealth, 2023. Source of the 349,114-epoch comparison and the 0.06 to 0.56 kappa range. https://pmc.ncbi.nlm.nih.gov/articles/PMC10654909/
- Lee YJ, Lee JY, Cho JH, Choi JH. Interrater reliability of sleep stage scoring: a meta-analysis. Journal of Clinical Sleep Medicine, 2022. Source of the human inter-scorer kappa figures. https://pmc.ncbi.nlm.nih.gov/articles/PMC8807917/
- ANSI/CTA/NSF-2052.1-A. Definitions and Characteristics for Wearable Sleep Monitors. https://www.thensf.org/wp-content/uploads/2022/10/ANSI-CTA-NSF-2052.1-A-FINAL.pdf
- ANSI/CTA/NSF-2052.2-A. Methodology of Measurements for Features in Sleep Tracking Consumer Technology Devices and Applications. https://www.thensf.org/wp-content/uploads/2025/03/ANSI-CTA-NSF-2052.2-A-FINAL.pdf
- Robbins R, Weaver MD, Sullivan JP, et al. Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors, 2024. Funded by Oura Ring Inc.; first author reports Oura Medical Advisory Board membership. https://pmc.ncbi.nlm.nih.gov/articles/PMC11511193/
- US Food and Drug Administration. General Wellness: Policy for Low Risk Devices, guidance for industry and FDA staff. https://www.fda.gov/media/90652/download
- Sahha. How Sahha scores are built, and what accuracy actually means. https://sahha.ai/guides/score-accuracy-explained/
- Sahha. Research study: design, data collection and ethical framework. University of Otago Human Ethics Committee reference 21/074. https://sahha.ai/research/research-study-protocol/