September 15, 2026 · 10 min read · Sugam Budhraja

What It Actually Takes to Validate a Health Score

Open source, expert-built, evidence-based and validated are four different claims, and the industry says them with the same word. What each one proves, what it does not, and the questions to ask any vendor including us.

Validating a health score means comparing its output against an external reference standard, in a study, using ground-truth labels the score did not have access to. Everything else that gets called validation is something weaker.

That distinction matters because a vendor telling you their scores are validated could mean any of four different things, and the gap between the weakest and the strongest version is the difference between a marketing decision and a research programme.

This is not pedantry. In August 2026, a proposed class action was filed against Oura over the gap between a number and the task that number describes. The allegations are unproven. But the underlying confusion is everywhere in this category, and it is worth being able to read a claim precisely, including ours.


Four claims, one word

LevelThe claimWhat it provesWhat it does not
1. OpenThe algorithm is publicExactly what it computesWhether the output corresponds to anything
2. Expert-builtA credentialed scientist designed itThe construct is not arbitraryThat the output was ever measured against reality
3. Evidence-basedEach component is grounded in published researchEvery factor and threshold is defensibleThat the composite was tested end to end
4. ValidatedOutput tested against a reference standard, in a studyThe output corresponds to something externalThat it generalises beyond that population and task

Most of this industry, including us on most scores, lives at level 3. Very little of it reaches level 4 for composite wellness scores. The problem is not that level 3 is bad. It is a genuinely strong position. The problem is that all four levels get described with the same vocabulary.


What open source proves

Publishing an algorithm is a real contribution and the recent move toward it in this category is good for buyers. Be clear about what it buys you.

It proves what the algorithm computes. You can read the thresholds, check the weighting, see whether the documentation matches the code, and confirm that nothing surprising happens to your users’ data on the way to a number. That is genuine, and it is more than most vendors offer.

It cannot prove the output means anything, because the thing that would prove that is not in the repository. Correctness here is not a property of code, it is a property of the relationship between the code’s output and the world, and you cannot read that relationship out of a source file. An open algorithm can be perfectly auditable and still be measuring the wrong construct. A closed one can be rigorously validated.

Openness and accuracy are close to orthogonal. Treating them as the same axis is the single most common error in reading these claims, and it is an easy one to make because transparency feels like honesty and honesty feels like correctness.

This cuts against a position we hold ourselves. Our own score accuracy guide argues that transparency, evidence, honesty about limits and consistency are a more durable answer than a single accuracy percentage. We still think that is right, and it is not a claim that transparency substitutes for validation. It is a claim that a visible factor breakdown you can interrogate is more useful to a product team than an unexplained number. Both things are true: transparency is worth a lot, and it is not measurement.

What “validated by a scientist” usually means

A credentialed researcher designing or reviewing an algorithm is worth having. It means thresholds came from the literature rather than from product intuition, and it is a meaningful step above assembling a score from guesses.

It is expert review, not measurement validation. The distinction is whether anyone compared the output to something external, with participants, a protocol and an analysis plan. A neuroscientist choosing a defensible HRV window is doing level 2 and possibly level 3 work. Neither becomes level 4 by virtue of the credential.

This matters because the formulation “designed and validated by [expert]” is common, reads as level 4, and usually is not. That is rarely deliberate. Validated is simply a word that has drifted.


Even real validation has a ceiling

The uncomfortable part, and the reason level 4 is harder than it sounds: the reference standard is not perfect either.

Polysomnography is the gold standard for sleep staging. A multicentre validation of eleven consumer trackers against PSG across 349,114 epochs found Cohen’s kappa for stage agreement ranging from 0.06 to 0.56, with no device reaching substantial agreement. That looks damning until you see the control: trained human scorers reading the same recording reach kappa around 0.57 on deep sleep and 0.24 on N1.

So the ceiling on agreement is not set by the wristband. It is set by the fact that two qualified humans reading identical data disagree substantially about what stage someone was in. A device cannot exceed the reproducibility of the thing it is being measured against.

Which is why a bare percentage is uninterpretable. Any validation number needs four things attached: the task (sleep versus wake, or four-stage classification), the metric (raw agreement, kappa, sensitivity and macro F1 give different numbers from identical data), the population (healthy young adults or a clinical cohort), and the reference (one scorer, a consensus, or an automated system).


The reporting standard almost nobody cites

For sleep specifically, a voluntary standard already exists. ANSI/CTA/NSF-2052, developed with the National Sleep Foundation, covers terminology for wearable sleep monitors and the methods for measuring their performance. It sets two separate compliance categories, one for sleep and wake determination and one for sleep stage determination, and asks that performance be reported as mean absolute difference together with a measure of its variability rather than as a single point estimate.

It is cited remarkably rarely. For composite wellness, readiness or resilience scores outside sleep, no comparable reporting standard exists at all, which is part of why the vocabulary drifted in the first place.


Where Sahha actually sits

Applying the ladder to ourselves, because a post like this is worthless otherwise.

The behavioural and mental health work reaches level 4. That rests on a longitudinal study rather than an argument: participants recruited through Prolific in waves from February 2022 with collection through mid-2025, passively monitored for five to twelve or more weeks, completing validated psychometric instruments weekly to provide ground-truth labels, under University of Otago Human Ethics Committee approval, reference 21/074. The study protocol is published, including its limitations. If you are comparing vendors, that reference number is the thing to ask every one of them for.

Score construction is level 3, which is where the strong version of this argument actually lives. Every factor and every goal in a Sahha score is grounded in published research, and our score accuracy guide sets this out factor by factor: seven to nine hours associated with lowest mortality risk, irregular sleep timing carrying risk independent of duration, circadian misalignment tied to metabolic and cardiovascular outcomes. A composite built from individually evidenced factors, each one visible and separately scored, is a defensible construct and a more useful one than an opaque number with a percentage attached.

What it is not is a claim that the composite sleep score has been validated against polysomnography, and we do not make that claim. Nor, as far as we can find, does anyone else in this category make it with a citation attached. Ground truth for depression, anxiety and stress does not transfer into a validated sleep-stage claim, and a company that ran one study does not get to describe its whole product as validated. We would rather draw that line ourselves than have a prospect find it.

The other thing worth saying: our scores return null rather than guessing when an input is missing, and phone-only factors are labelled separately from wearable-dependent ones. Returning nothing is a validity decision, not a gap in coverage.


Six questions for any vendor, including us

  1. What reference standard was the output compared against, if any? If the answer is none, that is fine, but it should be said rather than implied.
  2. Who were the participants, and how many? Thirty-five healthy adults for one night is a different claim from a year of longitudinal data.
  3. What metric is the number you quote, and for which task? A figure for sleep-versus-wake is not a figure for staging.
  4. What does the score return when an input is missing? A score that always produces a number is telling you less than one that sometimes refuses.
  5. Was there ethics approval and a pre-specified analysis? This separates a study from a retrospective look at data you already had.
  6. Can I see the factor breakdown behind a single score? If a number cannot be decomposed, no claim about it can be checked.

A vendor answering all six is being straight with you regardless of how flattering the answers are. A vendor answering none while using the word validated is worth pressing, and the pressing is cheap: every one of these is a question, not a procurement exercise. If you are running a wider evaluation, our comparison pages carry the dated, sourced version of what each platform publishes, and the science guides set out the factor-level evidence behind each Sahha score.

The honest summary of this category in 2026 is that almost nobody is at level 4 for composite scores, several are at level 3, a growing number are at level 1, and the marketing language does not reliably distinguish them. Knowing which rung a claim is standing on is most of the work.

References

  1. Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study. JMIR mHealth and uHealth, 2023. Source of the 349,114-epoch comparison and the 0.06 to 0.56 kappa range. https://pmc.ncbi.nlm.nih.gov/articles/PMC10654909/
  2. Lee YJ, Lee JY, Cho JH, Choi JH. Interrater reliability of sleep stage scoring: a meta-analysis. Journal of Clinical Sleep Medicine, 2022. Source of the human inter-scorer kappa figures. https://pmc.ncbi.nlm.nih.gov/articles/PMC8807917/
  3. ANSI/CTA/NSF-2052.1-A. Definitions and Characteristics for Wearable Sleep Monitors. https://www.thensf.org/wp-content/uploads/2022/10/ANSI-CTA-NSF-2052.1-A-FINAL.pdf
  4. ANSI/CTA/NSF-2052.2-A. Methodology of Measurements for Features in Sleep Tracking Consumer Technology Devices and Applications. https://www.thensf.org/wp-content/uploads/2025/03/ANSI-CTA-NSF-2052.2-A-FINAL.pdf
  5. Robbins R, Weaver MD, Sullivan JP, et al. Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors, 2024. Funded by Oura Ring Inc.; first author reports Oura Medical Advisory Board membership. https://pmc.ncbi.nlm.nih.gov/articles/PMC11511193/
  6. US Food and Drug Administration. General Wellness: Policy for Low Risk Devices, guidance for industry and FDA staff. https://www.fda.gov/media/90652/download
  7. Sahha. How Sahha scores are built, and what accuracy actually means. https://sahha.ai/guides/score-accuracy-explained/
  8. Sahha. Research study: design, data collection and ethical framework. University of Otago Human Ethics Committee reference 21/074. https://sahha.ai/research/research-study-protocol/

Related