August 24, 2026 · 13 min read · Sugam Budhraja

Where Oura's 95% Sleep Accuracy Claim Actually Comes From

The number is real and traceable to a peer-reviewed study. It is sleep-versus-wake sensitivity, not sleep staging, and in the same study Apple scored higher. Four-stage agreement in that study was 76.3%. One table, two numbers, and the gap between them is the whole dispute.

On 20 August a proposed class action was filed in the Northern District of California against Oura, alleging that its marketing overstated how accurately its rings identify sleep stages. The complaint singles out a claim of “95% Sleep Staging Accuracy compared to [a] clinical sleep lab” and argues that sleep happens in the brain rather than on a finger [1]. Oura disputes the allegations, says it stands behind its science and research, and notes that the ring is not a medical device or a substitute for a clinical sleep study [1]. Nothing has been decided, and a filed complaint is an allegation rather than a finding.

Most coverage has repeated the figure without saying where it came from. It appears to be traceable, and tracing it explains the entire dispute better than either side’s statement does.

Oura’s own science page points to a 2024 study in Sensors as validation of its sleep staging algorithm. It was conducted at the Brigham and Women’s Hospital Center for Clinical Investigation, funded by Oura, and put an Oura Ring Gen3, an Apple Watch Series 8 and a Fitbit Sense 2 against polysomnography across 35 healthy adults for a single night [2]. Its results include both of these:

MeasureOuraAppleFitbit
Sleep vs wake sensitivity95%97%95%
Four-stage agreement76.3%75.0%70.9%
Four-stage Cohen’s kappa0.650.600.55

One study, one table, and the number you quote depends entirely on which row you pick. Ninety-five percent is real, peer-reviewed, and describes telling sleep apart from wake. Sleep staging, the four-way split the complaint is actually about, is 76.3% in the same table. Apple scored higher on the row that produces the 95%, so it is not even a differentiator.

Read those numbers with their provenance attached. Beyond the funding, the study’s first author sits on Oura’s medical advisory board and reports consulting fees from the company [2], and 35 healthy adults for a single night is a small sample. None of that makes the figures wrong, and all of it is disclosed in the paper itself, which is more than the marketing page does. The independent multicentre data below is less flattering to every device in it, including this one.

So the fair reading is narrower than either side is arguing. The complaint is right that a number describing one task appears to have been labelled with the name of a different one. Oura is right that its accuracy work is real, published and independently conducted. And the disagreement is not about whether a finger can read a brain. It is about which row of a table gets quoted, and about the fact that nothing anywhere requires anyone to say.


The reference standard disagrees with itself

“Compared to a clinical sleep lab” sounds like a comparison against truth. It is not, and this is the part most coverage skips.

A meta-analysis of interrater reliability put trained human scorers against each other on the same polysomnography recordings [3]. Overall agreement was substantial, at Cohen’s kappa 0.76 (95% CI 0.71 to 0.81). Broken down by stage, it looks quite different.

StageInterrater kappa95% CI
Wake0.700.63 to 0.77
N1 (light)0.240.15 to 0.33
N20.570.54 to 0.60
N3 (deep)0.570.42 to 0.71
REM0.690.58 to 0.81

Two qualified experts reading the same night reach only moderate agreement on deep sleep, and only fair agreement on N1. The confidence interval on deep sleep runs from 0.42 to 0.71, so even the estimate of how much humans agree is loose.

A note on reading kappa, since the arithmetic matters here: it is chance-corrected, so 0.57 does not mean the scorers agreed 57% of the time. Raw agreement is higher. It means agreement is moderate once you remove what you would expect from guessing.

So a wearable validated “against a clinical sleep lab” is being scored against a panel of humans who disagree with each other about deep sleep. There is a ceiling, and it is not 100%.


Which is why the whole category’s numbers look the way they do

A multicentre validation ran eleven consumer trackers against polysomnography across 349,114 epochs from 75 participants at two Korean centres [4]. Cohen’s kappa for stage agreement ranged from 0.06 to 0.56, and no device in that study reached substantial agreement. The Apple Watch Series 8 scored 0.2976. The Fitbit Sense 2, among the stronger performers, scored 0.4185.

Hold that against the Brigham numbers above, where Oura reached 0.65 and Apple 0.60. Two peer-reviewed validations, the same device generation, and verdicts on opposite sides of the line between moderate and substantial agreement. The samples differ, the labs differ, the algorithm versions may differ, and one was funded by the manufacturer. All of which is the point: if the published literature cannot converge on a single number, no marketing page can honestly compress it into one.

Set those next to the human numbers and the picture is more interesting than “trackers are wrong.” The better devices are approaching the band where trained humans sit for the harder stages. Different studies and different populations, so this is a comparison of magnitude rather than a like-for-like scoreboard. But it does not describe a product category failing against a fixed truth. It describes a hard classification problem where nothing, including the reference method, performs cleanly.

The honest framing. Estimation is normal instrument science. A pulse oximeter estimates arterial oxygen saturation from light absorption. A cuff estimates arterial pressure from the sound of turbulent flow. Neither measures the quantity in its name directly, and both are clinically useful. Estimating sleep stages from heart rate, movement and temperature belongs in that tradition. Improving the estimate over time is how measurement technology has always progressed.

The problem is that “95%” is not a statement about anything

If estimation is legitimate, where does the exposure come from? From the fact that a single percentage can mean at least four different things, and the reader cannot tell which.

Which task. Sleep versus wake is an easy classification and devices do it well. Four-stage classification is a much harder one and devices do it far less well. A single accuracy figure that does not say which task it describes is not checkable.

Which metric. Raw agreement, kappa, sensitivity, specificity and macro F1 all produce different numbers from the same data. Raw agreement flatters, because most of the night is one or two stages.

Which population. Healthy young adults produce better numbers than older adults or people with sleep disorders, which is often the group most motivated to buy the device.

Against which reference. One scorer, a consensus of scorers, or an automated scoring system, each of which has its own error.

A number without those four attached cannot be verified or disproved. That is what makes it a marketing claim rather than a scientific one, and it is the thing a plaintiff’s lawyer reaches for first.


A disclosure standard already exists, and almost nobody cites it

This problem has already been solved on paper. A voluntary standard, written with the National Sleep Foundation, defines the terminology for wearable sleep monitors [5] and the methods for measuring how they perform [6].

Two things in them matter more than the rest.

First, they establish two separate compliance categories, one for sleep and wake determination and one for sleep stage determination. That single split would prevent most of the confusion above, because it forces a vendor to say which task a number describes.

Second, they ask that performance be reported as the mean of absolute differences against the reference together with a measure of its variability. A point estimate on its own is not compliant reporting. A number with a spread is.

The standard is voluntary, it has existed in some form since 2016, and you will struggle to find a consumer sleep product whose marketing page cites compliance with it. That, rather than the existence of estimation, is the actual gap.


So who is responsible for this

Worth separating, because the answers are different and only one of them is workable.

The FDA has deliberately stepped back. Its revised General Wellness guidance, issued in January 2026, brought noninvasive products that measure physiological parameters under enforcement discretion where they are genuinely intended for wellness use, reversing the more restrictive position behind earlier enforcement [7]. One of its worked examples is a wrist-worn device tracking sleep hours and sleep quality. This is the right call. You do not want a premarket submission gating a sleep score, and the agency has said as much.

The line it draws is about language, not sensors. Claims that reference a specific disease, or that diagnose, treat or mitigate, land you in device territory. Whoop learned the shape of this in public: the FDA’s July 2025 warning letter over Blood Pressure Insights was answered by changing the product and its labelling, and the agency closed the letter in June 2026 [8]. Half the fix was what the product said about itself.

The FTC covers the advertising, after the fact. Its Health Products Compliance Guidance asks for competent and reliable scientific evidence behind health-related claims [9]. That is the regime a “95% accurate” claim actually sits in. But it bites retrospectively and mostly on the loudest claims, so it sets no standard you can build against in advance.

Which leaves the plaintiffs’ bar as the default regulator, and that is what the Oura filing is. It is a poor mechanism for this. It arrives years late, it cannot distinguish a sloppy marketing page from a bad algorithm, and it produces settlements rather than standards. Nobody learns what good disclosure looks like from a docket.

The split that actually works is between the data provider and the developer. The provider owes error characteristics: which task, which metric, which population, which reference, with a spread. The developer owes presentation: what precision to render, what uncertainty to surface, and what the number is allowed to be used for inside the product. Neither party is required to do either today, which is precisely why the standard that would let them both do it goes uncited.


Claims you can defend

The pattern is that the exposed version names an outcome, and the defensible version names a method.

Instead ofWriteWhy
Measures your deep sleepEstimates time in deep sleep from heart rate, movement and temperatureThe verb should match the method. Nothing on a wrist measures brain activity.
95% accurateSleep and wake agreement of X against polysomnography in N adults; stage-level agreement of YNames the task, reference and population, so it can be checked
Clinically validatedValidated against [named instrument] in [named study], method published”Validated” without a nameable study is the claim with the least support and the most exposure
Compared to a clinical sleep labCompared against consensus polysomnography scoring, which itself reaches kappa 0.57 on deep sleepSets an honest ceiling rather than implying a perfect reference
Detects sleep apneaFlags irregular breathing patterns for discussion with a clinicianThe first is a disease claim and needs clearance. The second does not.
1 hour 42 minutes of deep sleepAround 1.5 hours, or a trend against the user’s own baselinePrecision is itself a claim, and the method does not have minute-level resolution

That last row is the one product teams underrate. Rendering a classifier output to the minute asserts a resolution the classifier does not possess, and no disclaimer elsewhere in the app undoes it. The interface is a claim, whatever the legal page says. Showing a range, a confidence band, or a movement against the user’s own baseline costs very little and is far easier to defend than a number that looks measured.

Trends are also the safer ground regulatorily. Reporting deviations from a user’s personal baseline, in language that avoids disease terms, is squarely inside what the January 2026 guidance describes as wellness.


The rule we hold ourselves to

All of this is easier to write than to follow, so here is where we land on our own pages.

We use “validated” only where we can name what the validation points at. Our mental wellbeing models are trained and evaluated against clinically used instruments, principally PHQ-9 and DASS-21, and we publish the study design, dataset and method. That page also says plainly that the score is not a diagnostic instrument and does not carry peer-reviewed clinical validation of its own. The decades of validation belong to the reference instruments it is trained against, which is a different sentence and the accurate one.

For the sleep score we say “research-backed,” because that is what it is. It is built from seven factors with literature behind each one, which is a claim about how it is constructed rather than a claim about a validated composite output. Those are two different assertions and they deserve two different words.

The wider point for anyone writing this copy: drift happens because marketing pages and research programmes move at different speeds, not because anyone sets out to overstate. Which is why it is worth reading your own pages with the four questions above in hand, on a schedule, rather than waiting for someone else to do it for you.

Not legal advice. This is a developer’s reading of the public record, not counsel. If you are shipping health claims at scale, have a lawyer read your marketing page, and read what a phone can and cannot detect about sleep stages for the sensor-level detail behind the numbers above.

References

  1. TechCrunch. Oura faces lawsuit accusing it of misleading consumers about sleep-tracking accuracy, August 2026. https://techcrunch.com/2026/08/21/oura-faces-lawsuit-accusing-it-of-misleading-consumers-about-sleep-tracking-accuracy/
  2. Robbins R, Weaver MD, Sullivan JP, Quan SF, Gilmore K, Shaw S, Benz A, Qadri S, Barger LK, Czeisler CA, Duffy JF. Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors, 2024. Funded by Oura Ring Inc.; first author reports Oura Medical Advisory Board membership and consulting fees. https://pmc.ncbi.nlm.nih.gov/articles/PMC11511193/
  3. Lee YJ, Lee JY, Cho JH, Choi JH. Interrater reliability of sleep stage scoring: a meta-analysis. Journal of Clinical Sleep Medicine, 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC8807917/
  4. Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study. JMIR mHealth and uHealth, 2023. https://pmc.ncbi.nlm.nih.gov/articles/PMC10654909/
  5. ANSI/CTA/NSF-2052.1-A. Definitions and Characteristics for Wearable Sleep Monitors. https://www.thensf.org/wp-content/uploads/2022/10/ANSI-CTA-NSF-2052.1-A-FINAL.pdf
  6. ANSI/CTA/NSF-2052.2-A. Methodology of Measurements for Features in Sleep Tracking Consumer Technology Devices and Applications. https://www.thensf.org/wp-content/uploads/2025/03/ANSI-CTA-NSF-2052.2-A-FINAL.pdf
  7. US Food and Drug Administration. General Wellness: Policy for Low Risk Devices, guidance for industry and FDA staff. https://www.fda.gov/media/90652/download
  8. US Food and Drug Administration. Warning Letter, WHOOP Inc., 709755, 14 July 2025. https://www.fda.gov/inspections-compliance-enforcement-and-criminal-investigations/warning-letters/whoop-inc-709755-07142025
  9. US Federal Trade Commission. Health Products Compliance Guidance, December 2022. https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance

Related