August 31, 2026 · 11 min read · Sugam Budhraja

Can a Watch Detect Insulin Resistance? Google's Own Study Says It Misses 40% of Cases

Google shipped Health Guardian with Insulin Resistance Trends and described its models as rigorously validated against gold-standard clinical measurements. The underlying paper is good, public, and says something more specific: the gold standard is the euglycemic clamp, the study used HOMA-IR instead, and the wearable-only model reaches sensitivity 0.60. The gap is between the paper and the product page.

On 12 August 2026, Google announced the Pixel Watch 5 and a set of features called Health Guardian: monthly trend summaries for blood pressure, sleep breathing quality, and insulin resistance, with the first report arriving after a month of wear [7]. The launch post says the underlying health foundation models were “trained on billions of minutes of sensor data from opted-in users and rigorously validated against gold-standard clinical measurements” [1].

Estimating insulin resistance from a wrist is a real research result, and Google has published the work rather than asserting it. That is more than most of this category does.

It is also the same shape as Oura’s 95% figure: the study is sound, and the sentence on the product page describes something the study did not do.


What the study actually is

The relevant work is WEAR-ME, from Google Research with the University of Cambridge Institute of Metabolic Science, published in Nature [2] [3] [4]. It enrolled 1,165 participants remotely through the Google Health Studies app: median BMI 28, median age 45, median HbA1c 5.4%.

Participants contributed wearable data from a Fitbit or Pixel Watch, plus blood work drawn at Quest Diagnostics, plus demographics. Insulin resistance was defined as HOMA-IR above 2.9, giving a cohort of 459 insulin sensitive, 406 impaired, and 300 insulin resistant, so about 26% prevalence.

The wearable inputs are deliberately modest: resting heart rate, daily step count, sleep duration and heart rate variability, used as aggregate measures rather than time series. The authors say this was intentional so the models would run on lower-cost devices [2]. The features most correlated with HOMA-IR were resting heart rate, daily steps and HRV.


The number depends entirely on the configuration

This is the part that matters, and it is stated plainly in the paper.

Model inputsauROCSensitivitySpecificity
Wearables and demographics0.700.600.80
Wearables, demographics, fasting glucose0.780.730.84
Wearables, demographics, lipid and metabolic panel0.800.760.84

The 0.80 that gets quoted in coverage is the bottom row. It requires a blood draw, a lipid panel and a metabolic panel.

The row that describes what a watch on a wrist can actually see, with no blood test, is the top one. auROC 0.70. Sensitivity 0.60.

A sensitivity of 0.60 means that of every 100 people who are insulin resistant by the study’s own definition, the wearable-only model identifies 60 and misses 40.


What sensitivity 0.60 does to a consumer feature

Sensitivity and specificity are properties of the model. What a user experiences is positive predictive value, which depends on how common the condition is among the people wearing the device.

At the wearable-only figures, sensitivity 0.60 and specificity 0.80:

Prevalence of insulin resistanceA flag is correctAn all-clear is correct
10%, a healthier buyer population25%95%
15%35%92%
25.8%, the WEAR-ME cohort51%85%
40%, a high-risk population67%75%

Put concretely, at 15% prevalence, per 1,000 people wearing the watch: 260 get flagged, of whom 170 are not insulin resistant. A further 60 are insulin resistant and get told nothing.

This is not a criticism of the model. It is what a 0.70 auROC does when it meets a low-prevalence population. Every screening test behaves this way, which is why clinical screening programmes are targeted rather than universal. The issue is that a consumer wearable is the most untargeted deployment surface that exists: it goes to whoever buys the watch, and the marketing does not say that the feature works well for some of them and poorly for others.

The reference standard is not the gold standard

The launch language is “rigorously validated against gold-standard clinical measurements” [1]. The paper is more precise, and contradicts the implication in its own words.

It identifies the gold standard directly: “The gold standard test for IR is the hyperinsulinemic euglycemic clamp,” noting it “is performed in research facilities only, expensive, and time-consuming” [2] [5]. It then explains that HOMA-IR and similar measures “are affordable and faster alternatives that can be performed in clinical labs,” and adds that they “may not be as accurate as the clamp” [2].

The study validated against HOMA-IR, not the clamp.

It goes further and names the noise in its own reference. Using HOMA-IR as ground truth “presents some challenges, primarily related to per-subject reproducibility and standardization of insulin concentrations across laboratories. This can lead to a reported coefficient of variation of 23.5% between two measurements” [2] [6].

So the same person, measured twice, can produce HOMA-IR values 23.5% apart. That is the yardstick the model was fitted and scored against. The authors handle this honestly, standardising all draws through Quest to reduce cross-lab variation, and concluding that HOMA-IR “remains a valuable tool for classifying individuals into broad categories of insulin sensitivity or resistance, rather than providing precise quantifications” [2].

That sentence is the correct claim. “Rigorously validated against gold-standard clinical measurements” is a different claim, and the paper it points at does not support it.


What Google got right

A post that stopped there would be unfair, and would miss the most interesting finding.

The subgroup result is genuinely strong. The paper reports 93% sensitivity and 95% adjusted specificity in obese and sedentary participants, described as the subpopulation most vulnerable to developing type 2 diabetes and most likely to benefit from early intervention [2]. The model performs best exactly where the clinical value is highest. That is the opposite of the usual failure mode, where a model looks good on average and collapses on the cases that matter.

The work is public. There is a paper, an independent validation cohort of 72, named authors, and a stated reference standard with its limitations spelled out. Almost nothing else in consumer health scoring offers that, which is the whole argument of our piece on what validation actually means.

And the underlying finding is real. Resting heart rate, steps, sleep and HRV carry enough signal about metabolic health to beat demographics alone by a wide margin. That is a legitimate result and it will get better.

The problem is one sentence on a product page, not the science under it.


Notice what the feature is not called. Not Insulin Resistance Score, not Insulin Resistance Detection, not a HOMA-IR estimate.

The FDA’s revised General Wellness guidance, issued 6 January 2026, added a passage on exactly this class of product: non-invasive sensing used “to estimate, infer, or output physiologic parameters,” with blood glucose named among the examples. Products meeting its conditions, the guidance says, “may display values, ranges, trends, baselines, or longitudinal summaries, and may contextualize these outputs in relation to sleep, activity, stress, recovery, or similar wellness domains” [8].

So a monthly trend summary is not a softer version of a real feature. It is the shape the guidance explicitly names as permitted. The naming is not marketing timidity, it is the product built to sit where the January guidance says it can sit. We went through the full conditions in where the FDA line sits now.

Two things in that same guidance are worth holding against this product, though, and neither is settled.

The exclusion list names screening. Products are not general wellness products, the guidance states, “when they are intended to measure, estimate, or report physiologic values for medical or clinical purposes, including screening, diagnosis, monitoring, alerting, or management of a disease or condition” [8]. Insulin resistance is a clinical condition, and a feature that tells a general population whether they may be developing it is doing something adjacent to screening. Framing it as a wellness trend rather than a result is what keeps distance from that word.

And the guidance attaches a validation condition. The last of its six conditions is that a product “do[es] not include values that mimic those used clinically unless validated (e.g. manufacturer testing, peer-reviewed clinical literature) to reflect those values” [8]. Both worked examples in the guidance repeat it, qualifying wellness status on the product having “validated values.”

That makes the marketing sentence load-bearing rather than decorative. “Rigorously validated against gold-standard clinical measurements” is not only a claim to consumers. Validation is a condition the guidance itself attaches to this exact category of product. Which makes it more, not less, important that the published study validated against HOMA-IR while identifying the euglycemic clamp as the gold standard, and that no device-level validation for the shipping feature has been published. Whether the condition is met is a question for Google and its regulators, not for us. That it is a question at all is the point.

Which is also worth knowing if you are building something similar, because the regulatory design and the honest design converge. A reference standard with 23.5% intra-person variability and a model at 0.60 sensitivity cannot support a number. It can support a direction.


What to take from this if you are shipping something like it

Publish the configuration, not just the headline. If your model performs differently with and without an input your product does not have, the number that belongs on the page is the one matching what ships. Quoting the best row of a table is the single most common way an honest study becomes a misleading claim.

Say who it works for. A model at 93% sensitivity in high-risk users and 60% overall is two different products. Users in the low-risk majority are being handed the weak version with the strong version’s marketing.

Do not use “gold standard” unless you used the gold standard. It is a term of art with a specific referent in most clinical areas, and the paper you are citing will usually name it. If you validated against a convenient proxy, name the proxy and its error.

Present direction, not magnitude, when your reference is noisy. This is the recurring conclusion across everything we have looked at, from sleep staging to step counts: relative change within one source is far more defensible than an absolute value, and it is usually more useful to the user anyway.


The short version

Google shipped a wearable feature that estimates insulin resistance, backed by a real published study of 1,165 participants that most of this category would not have produced.

That study reports auROC 0.80 with a lipid and metabolic panel, and auROC 0.70 with sensitivity 0.60 from wearables and demographics alone, which is the configuration a watch actually has. At realistic consumer prevalence, most flags from that model are false positives and 40% of genuine cases are missed. It was validated against HOMA-IR, which the paper itself identifies as not the gold standard and as varying 23.5% between two measurements of the same person.

The paper says all of this clearly. The product page says “rigorously validated against gold-standard clinical measurements.”

Fairness note. Nothing here alleges the feature does not work, and the model is strongest in exactly the population that benefits most. Google has not published performance figures for the shipping feature specifically, so the numbers above come from the published study that the feature is built on, and the shipping model, SensorFM, is a later iteration than the one evaluated in that paper. If Google publishes device-level validation, this post should be read against that instead.

References

  1. Pixel Watch 5 is here with new health features and Gemini AI. Google, 12 August 2026. Retrieved 31 August 2026. https://blog.google/products-and-platforms/devices/pixel/pixel-watch-5/
  2. Metwally AA, Heydari AA, McDuff D, Solot A, Esmaeilpour Z, Faranesh AZ, Zhou M, Savage DB, Heneghan C, Patel S, Speed C, Prieto JL. Insulin Resistance Prediction From Wearables and Routine Blood Biomarkers. Google Research and University of Cambridge Institute of Metabolic Science. arXiv:2505.03784. https://arxiv.org/abs/2505.03784
  3. Insulin resistance prediction from wearables and routine blood biomarkers. Google Research blog. Retrieved 31 August 2026. https://research.google/blog/insulin-resistance-prediction-from-wearables-and-routine-blood-biomarkers/
  4. Insulin resistance prediction from wearables and routine blood biomarkers. Nature. https://www.nature.com/articles/s41586-026-10179-2
  5. DeFronzo RA, Tobin JD, Andres R. Glucose clamp technique: a method for quantifying insulin secretion and resistance, 1979. Cited in [2] as the gold standard method.
  6. Sarafidis PA et al., 2007. Source of the 23.5% coefficient of variation between two HOMA-IR measurements, as cited in [2].
  7. Google Health Guardian availability and feature description, August 2026. Engadget. https://www.engadget.com/2235228/google-health-guardian-pixel-watch-5-blood-pressure-insulin-resistance/
  8. US Food and Drug Administration, Center for Devices and Radiological Health. General Wellness: Policy for Low Risk Devices. Issued 6 January 2026. Quotations taken from the issued PDF. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/general-wellness-policy-low-risk-devices

Related