Abstract
Most people who use health apps do not wear a sleep tracker, but nearly all of them carry a phone. We measure how well sleep timing can be recovered from a phone alone, passively, using only what the phone already records: its step counter, screen locks and app events, with no location, microphone, camera or content, and no nightly action from the person. We describe an estimator built on an overnight quiet window, and we measure what each of its design choices contributes. Two observations drive the design. First, a wearable’s steps leak into its owner’s phone records, and settings developed on that data are wrong for phone-only users. Second, the time from the last phone activity to sleep is not a constant: it grows by about 50 minutes for every hour the phone goes quiet before the person’s usual time. Under an analysis plan fixed before evaluation, we validate the estimator against consumer wearables on 105,826 nights from 3,868 people held out from all development. It produces a measured estimate on 90% of nights with any phone activity (83% of all nights), with median errors of 24 minutes for sleep onset and 19 minutes for wake. Onset and wake are both within an hour of the wearable on 66% of nights, against 44% for the previous version and 75% for two wearables worn on the same night. Weekly averages of sleep timing are within 15 minutes, and the third of nights labelled high-confidence are within an hour 81% of the time. Single-night duration remains the weakest output (ICC 0.45).
Key findings
- 66% of nights within an hour, on onset and wake together, against 44% for the previous version. Two wearables worn on the same night manage 75%. A typical night is 19 to 24 minutes off at each end.
- Nine in ten nights with any phone activity get a measured estimate (83% of all nights). The rest are filled from weaker evidence and labelled as such.
- Weekly averages are sharp. A seven-night average of sleep timing is within 15 minutes of the wearable.
- Sleep starts later the earlier the phone goes quiet: 31 minutes at the usual time, about 50 minutes more for every hour early.
- The confidence label means what it says. High-confidence nights, a third of all nights, are right 81% of the time; low-confidence nights 39%.
- Wearable data quietly misleads phone-only development. Settings tuned on it cost the previous version 9 points.
1. Introduction
Sleep features in health, wellbeing and insurance apps usually depend on a wearable. Most users do not have one, or do not wear it to bed. In the four months covered here, about three in four active users of apps built on Sahha had no wearable sleep data. For them, a sleep score, a regularity measure or a bedtime nudge needs another source, and the phone is the only one present.
Inferring sleep from phone use is an established idea [5, 6]. Earlier work showed that phones capture the timing of sleep much better than its duration [7], and that time of day alone is a strong baseline [8]. It also leaves open the questions that decide whether such an estimator can be deployed. Cohorts are small, from a handful of people to a few hundred. References are diaries or research actigraphs. Results are often averaged per person rather than reported per night. Coverage, failure rates and confidence are rarely measured. And some methods rely on sensors that raise privacy concerns, such as the microphone [5].
We ask how far a strictly passive phone signal can go, and what it takes to get there. Our estimator finds the overnight stretch in which the phone is quiet and places sleep inside it (figure 2). The evidence for each component is the substance of the paper. We make four contributions:
- A finding about sleep latency. The time from the last phone activity to sleep onset depends on how early that activity ends, both by the clock and relative to the person’s habits (section 4). Modelling this does most of the work of the final estimator.
- A pitfall in phone-only development. A wearable’s steps leak into its owner’s phone records. Settings tuned on such data silently fail phone-only users; removing one such setting gains 9 points (section 8.2).
- A large held-out validation. 105,826 nights from 3,868 people, scored against consumer wearables with the standard framework for sleep trackers [3], adding coverage, tails, calibrated confidence and the agreement of weekly averages, run once under a plan fixed in advance (section 7).
- An account of what matters. An ablation of every design choice, from which signals to use to how sleep is placed, on held-out nights (section 8).
2. Related work
Measuring sleep outside the laboratory. Polysomnography is the reference for sleep staging but is confined to a laboratory or a single night at home. Wrist actigraphy infers sleep from movement and has long served as the field alternative. Consumer wearables add heart rate and its variability, and now estimate sleep for millions of people every night [1]. Against polysomnography they detect sleep well and wake less well, and they tend to overestimate total sleep by tens of minutes, with performance varying by device and by person [1, 2]. A standardised framework for testing sleep trackers recommends reporting discrepancy alongside minute-by-minute agreement [3]; we follow it. Consumer devices now support very large observational studies, such as five million nights from a smart ring [4].
Sleep from phone use. The Best Effort Sleep model combined phone use, charging, silence and darkness, and estimated duration within about 42 minutes over an eight-person week [5]. Screen-unlock rhythms reflect chronotype and sleep debt (nine people over 97 days) [6]. iSenseSleep used screen events alone, with an average duration error of 24 minutes in a small validation, rising to 68 to 83 minutes in some groups [9]. SensibleSleep fitted a Bayesian rest model to screen events and reported 89% accuracy [10]. Touchscreen interactions in 79 people over about 1,400 nights tracked actigraphy onset and wake closely (R² 0.84 and 0.90) but not duration (R² 0.28) [7]. Later work extended this to typing dynamics [11].
Sleep from phone sensors. Combining phone sensors against self-report gave median absolute deviations of 58 minutes for duration and 43 and 38 minutes for onset and wake; time of day alone classified sleep with 86.9% accuracy, against 88.8% for all sensors [8]. Studies in students using accelerometer and light data report good person-level agreement with diaries but rarely per-night error [12, 13]. Passive Wi-Fi sensing reaches similar timing accuracy at campus scale [14]. Phone use has also been combined with wearables to characterise sleep behaviour rather than replace them [15].
Why coverage and timing matter. In a 2019 survey, 21% of US adults regularly wore a smart watch or fitness tracker: 31% in households earning $75,000 or more, against 12% below $30,000 [23]. By 2025, 91% owned a smartphone [24]. A programme that scores sleep only from wearables reaches a minority of its members, and not a random one. Among the measures a phone can support, timing carries real weight: sleep regularity predicted all-cause mortality more strongly than sleep duration in 60,977 UK Biobank participants [16], irregular sleep is associated with later circadian timing and poorer outcomes [17], and social jetlag captures the conflict between body clock and social schedule [18]. Phone apps have already been used to map sleep timing across countries [19].
Positioning. No prior study, to our knowledge, evaluates a phone-only estimator against consumer wearables at this scale, on held-out people, under a plan fixed in advance, while reporting coverage, failure and calibrated confidence.
3. Problem setting
3.1 Task
A night runs from 18:00 to 14:00 the next day, in the person’s local time. For each night, the phone supplies a sequence of timestamped activity records, and a wearable, where one is worn, supplies a reference sleep onset and wake. The estimator sees only the phone’s records, from this night and earlier ones. It must output a sleep onset, a wake time, a sleep duration and a confidence label, or, when the night has no quiet window, an estimate from weaker evidence. We call the first a measured estimate and the second a filled one.
3.2 Data and privacy
Data came from end users of apps built on Sahha’s platform between 2 June and 30 September 2026. Identifiers were replaced with pseudonyms before analysis. The analysis used activity, device-event and sleep records and the type of device that recorded them; it used no location, audio, images, typed text, message content or names. No identifier appears in any result. Sahha carried out this analysis to validate and improve its own services.
3.3 The wearable reference
Reference sleep came from consumer wearables (watches, rings and one band) whose data reached the phone through the platform’s health store. Third-party apps that infer sleep from the phone or from another device’s data, bed sensors, relays between health platforms and manual entries were not used. Nights referenced by certain smart rings were withheld from model development and serve as a test on an unseen device (section 7.7). A night’s reference is the device’s main sleep episode between 18:00 and 14:00: asleep, light, deep and REM intervals are merged, runs less than 90 minutes apart are joined, and the episode with the most sleep is kept. Onset and wake are its bounds; total sleep time is the asleep time within it.
3.4 Phone-only, strictly
A wearable owner’s phone also receives the wearable’s steps. On 42% of wearable nights, something other than the phone recorded activity between midnight and 05:00. An estimator developed or tested on that data would use information a phone-only user never has (section 8.2). Every analysis here therefore uses only phone-origin records, identified by the recorder that wrote them, not by the device type attached to the record.
3.5 Metrics
Errors are estimate minus reference, so positive means later or longer. Our primary measure is the share of nights on which onset and wake are both within 60 minutes of the reference: a night counts only if the whole sleep period is placed correctly. We also report median absolute errors, the share within 30 minutes, and the share with onset or wake more than two hours off. Following the standard framework [3], we add discrepancy (bias, 95% limits of agreement, proportional bias, intraclass correlation) and minute-by-minute agreement (sensitivity for sleep, specificity for wake, accuracy, Cohen’s κ). Because a phone estimator must be judged on what it misses, we report coverage, calibration of the confidence label, and agreement of weekly averages and within-person trends. Intervals are 95% percentile intervals from 1,000 bootstrap resamples of people.
4. Two observations
Two patterns in the data shape the method. Both were found on development data; the figures here show them on held-out people who were never used to find them.
4.1 The phone goes quiet when sleep begins
Aligned to the wearable, phone use is steady through the evening and stops as sleep begins; almost nothing happens while the person is asleep; and use returns at once in the morning, with the first pickup a median of 18 minutes after the wearable records waking (figure 3). An overnight stretch with no phone activity, which we call the quiet window, therefore brackets the night. Its edges are close to sleep, but not on it.
4.2 The earlier the phone goes quiet, the longer until sleep
The gap between the last phone activity and sleep onset is not constant. When the phone goes quiet at the person’s usual time, sleep follows a median of 31 minutes later. Each hour earlier adds about 50 minutes: 76 minutes at one to one and a half hours early, 153 at two to three hours and 202 at three to four. When the phone goes quiet later than usual, sleep follows within 15 to 22 minutes.
Clock time has an effect of its own (figure 4). At the person’s usual time, the wait is 56 minutes when the phone goes quiet before 21:00 and 20 minutes after midnight. Habit matters at a fixed clock time: with the phone going quiet between 22:00 and 23:00, the median wait is 24 minutes for people who usually go quiet before 22:30 and 82 minutes for those who usually go quiet after 00:30.
Both effects fit the circadian timing of sleep. The hours before habitual bedtime, the wake maintenance zone or “forbidden zone” for sleep, are the hardest time to fall asleep [20, 21], and an early evening is more likely to fall inside them. Phone use in bed is also associated with later sleep onset [22]. Some of the longer waits will be time spent reading or watching television without the phone; the estimator does not need to tell these apart, because either way the person is not yet asleep.
5. The estimator
The estimator, v2, runs in five steps (figure 2a).
Inputs. v2 uses three phone-origin signals: the phone’s own step counter, screen lock and unlock, and the app opening, closing or resuming. These are signals the phone records anyway. v2 does not use location, the microphone, the camera, ambient light, Wi-Fi, typing, or the content of anything the person does on the phone. It asks nothing of the person: no diary, no button to press, nothing to wear or charge. The step counter needs a one-time motion permission on iPhone and Android; after that, it needs no action from the person. Activity recorded by wearables and other apps is ignored.
Quiet window. Activity isolated by more than 90 minutes on either side is dropped, so that a single glance at the phone does not split the night. The window is then the first gap of 5 to 12 hours that starts between 18:00 and 04:00. When no gap qualifies, a fallback window spans the last activity before 02:00 and the first after 04:00.
Placement. Following section 4, sleep is placed inside the window by two learned models, one for the delay after the window starts and one for the advance before it ends. Their inputs describe the window and its edges: when it starts and ends and how long it is; how much phone activity, screen locking and walking there was in the half hour to two hours either side; whether anything happened inside it; which kind of event closed and reopened it; the night’s total activity, the day of the week and the phone’s platform; and the person’s usual start and end over their previous 14 nights, where they exist.
Nights without a window. When neither rule finds a window, two further models estimate onset and wake from the evening’s last and the morning’s first phone activity, the longest gap in the night, the amount of overnight activity and the person’s usual times. These nights are labelled filled.
Duration and confidence. The phone cannot see wakefulness inside the night, so the window overstates sleep. A fifth model predicts duration from the placed window and the night’s activity. A sixth estimates the probability that onset and wake are both within an hour of the reference; it is reported as a confidence label (low below 0.5, medium from 0.5 to 0.75, high from 0.75).
All six models were fitted on development data from 30% of people and selected by five-fold cross-validation over people. Their inputs are described here; their fitted parameters are not published. Together they take a fraction of a millisecond per night to run.
6. Evaluation protocol
Split. People, not nights, were split at random: 30% for development and 70% held out, so no person contributes to both.
Lock. Every design choice was made on development data. The analysis plan (population, exclusions, reference definition, metrics, comparators, subgroups and figures) was then fixed and recorded with fingerprints of the data snapshot and the final models, and the evaluation in section 7 was run once. Two candidate estimators were scored in that run: v2 and a much smaller learned model. v2 was chosen for deployment afterwards, once its running cost had been measured; every number reported for it comes from the same single run. The ablations in section 8 were run on the held-out set afterwards, to explain choices already made; none changed the estimator.
Eligibility. Held-out people contributed 199,647 wearable nights. After excluding nights with two or more wearables (kept for section 7.6), time-zone changes, implausible reference sleep, nights used in an earlier calibration of v1 and nights with no phone-origin activity at all, 152,931 nights from 4,778 people were eligible (table A1). Setting aside the unseen-device nights (section 3.3) leaves the main evaluation set: 105,826 nights from 3,868 people.
Comparators. On the same nights, v2 is compared with v1 (the version it replaces), with a much smaller learned model, with each person’s median schedule over their previous 14 nights, with the population’s median schedule, and with a fixed 23:00 to 07:00 night.
7. Results
7.1 Accuracy on held-out people
On measured nights, v2 has onset and wake both within an hour of the wearable on 65.6% of nights (95% interval 64.7 to 66.4), against 44.2% for v1 (table 1). Its median errors are 24.2 minutes for onset and 19.2 for wake, and only 13.5% of nights are more than two hours off, against 31.1% for v1. Counting filled nights too, 62.2% of all nights are within an hour. v2 is essentially unbiased in timing, with median signed errors under one minute, where v1 started sleep 18 minutes too early and ended it 27 minutes too late. In plain terms, on a typical night the phone places sleep onset 24 minutes from the wearable and wake 19 minutes from it. Figure 1 shows random nights and the full distribution of errors.
| Measured nights (95,642 nights, 3,751 people) | v2 | 95% interval | v1 |
|---|---|---|---|
| Median absolute error, onset | 24.2 min | 23.7 to 24.8 | 32.5 min |
| Median absolute error, wake | 19.2 min | 18.8 to 19.7 | 33.4 min |
| Median absolute error, midpoint | 19.1 min | 27.4 min | |
| Onset and wake both within 30 min | 39.0% | 24.1% | |
| Onset and wake both within 60 min | 65.6% | 64.7 to 66.4 | 44.2% |
| Onset or wake more than 120 min off | 13.5% | 13.0 to 14.0 | 31.1% |
| Median signed error, onset / wake | +0.4 / -0.7 min | -18.5 / +27.0 min | |
| Filled nights: both within 60 min | 30.6% | 28.8 to 32.4 | |
| All nights, filled included: both within 60 min | 62.2% | 61.2 to 63.1 |
Table 1. Per-night agreement with the wearable. v1 figures are on the 91,315 nights it covers.
Against the alternatives. On nights both cover, v2’s per-night timing error (the mean of the onset and wake errors) is 13.8 minutes lower than v1’s (median paired difference; 95% interval 13.2 to 14.4) and 21.4 minutes lower than the person’s own recent schedule (20.9 to 21.9). The phone therefore tracks the night in question, not only the person’s habits. In the share of nights within an hour, v2 leads v1 by 21.9 points, the person’s own recent schedule by 27.6, the population’s schedule by 33.1 and a fixed 23:00 to 07:00 night by 36.0. A much smaller learned model is 4.6 points behind v2 (section 8.5).
Minute by minute. Scored minute by minute from 20:00 to 12:00, v2 agrees with the wearable on 89.2% of minutes, with sensitivity 0.923, specificity 0.868 and Cohen’s κ 0.781 (table A3). Time of day alone scores 80% accuracy, which is why accuracy figures for phone methods need a baseline beside them [8]; κ separates the methods more clearly (0.609 for a fixed 23:00 to 07:00 night, 0.692 for v1). v1’s high sensitivity and low specificity come from reporting the whole quiet window as sleep. v2’s estimate rises and falls at nearly the same times as the wearable, but stays higher in the small hours, because a phone cannot see brief wakefulness inside the window.
7.2 Coverage
Of 105,826 evaluation nights, 90.4% had a quiet window and received a measured estimate, and 9.6% were filled; v1 produced an estimate on 86.3%. These shares count nights on which the phone recorded at least one event. On a further 8,872 nights (7.7% of 114,698), the phone recorded nothing at all (for example, the phone was off or the app was not running); counting those, 83% of all nights receive a measured estimate.
Coverage needs no reference, so it can also be measured on people without a wearable (table 2, same rule). For people with only a phone, the group the estimator exists for, four in five of their nights with phone activity get a measured estimate, and the rest are filled.
| User group (held-out people) | Nights | v2 measured | v2 filled | v1 |
|---|---|---|---|---|
| Phone only | 41,471 | 79.0% | 21.0% | 74.2% |
| Phone and a daytime wearable without sleep tracking | 26,293 | 80.7% | 19.3% | 90.6% |
| Wearable sleep tracker | 417,732 | 87.5% | 12.5% | 86.8% |
Table 2. Coverage by user group, on nights with any phone activity. v1 also reads the steps of a wearable worn by day, which v2 deliberately ignores; that is why v1 covers more nights in the second group.
7.3 Confidence
The confidence label is calibrated: in every tenth of the predicted range, the share of held-out nights within an hour is within three percentage points of the prediction. High-confidence nights, 32% of all nights, have onset and wake both within an hour 81% of the time; medium-confidence nights, 41%, 65% of the time; low-confidence nights, 17%, 39%; and filled nights, 10%, 31% (figure 5b). Ranked by confidence, the half of nights v2 is most sure about are within an hour on 77%, and the most confident tenth on 85% (figure 5a).
A product can use the label directly: show high- and medium-confidence nights as measured, soften low-confidence nights, and present filled nights as estimates from weaker evidence.
7.4 Agreement and duration
Onset and wake errors are centred near zero with no visible drift across the night (figure 6). Mean differences are -7 minutes for onset and +6 for wake, larger than the medians in table 1 because a long tail of nights pulls the mean; 95% limits of agreement are -150 to +135 minutes for onset and -139 to +151 for wake, and intraclass correlations are 0.70 and 0.71. The sleep midpoint agrees best: bias -0.3 minutes, limits of ±120 minutes, intraclass correlation 0.75.
Mean duration bias is +4 minutes, against +117 minutes for v1. Single-night duration is nonetheless the weakest output: limits of agreement run from -139 to +148 minutes, intraclass correlation is 0.45, and a negative proportional bias shows that estimates are pulled toward typical durations. Averaging over a week brings the median duration error to 27 minutes (section 7.5).
7.5 Weekly averages and trends
Most uses of sleep data look at weeks, and averaging helps the phone because much of its nightly error does not repeat. On measured nights, the phone’s sleep midpoint correlates with the wearable’s at r = 0.77 for single nights and r = 0.91 for seven-night averages, where the median error is 13 minutes (figure 7). Over all nights, including filled ones, the median error of the midpoint falls from 21 minutes for a single night to 15 for a week and 13 for two weeks; for onset and wake it falls from 26 and 21 minutes to 20 and 18 for a week, and for duration from 40 minutes to 27 and 24.
Within a person, week-to-week changes in the phone’s average midpoint correlate with the wearable’s at r = 0.69 (637 people with at least eight weeks), and in average duration at r = 0.47. Weekly regularity, the standard deviation of sleep midpoint, correlates at r = 0.61, with a median error of 11 minutes. Social jetlag, the weekend minus weekday midpoint, correlates at r = 0.66, with a median error of 15 minutes over 1,966 people; the phone measures a median social jetlag of 36 minutes against the wearable’s 39.
7.6 How good is the reference?
On 5,155 nights where two wearables recorded the same person (343 people), they agreed on onset within a median of 5.5 minutes and on wake within 3.3 minutes; the phone was 28 and 22 minutes from them on those nights. Consumer wearables have their own error against polysomnography [2], but their agreement with each other shows that the phone’s error is mostly its own. Their disagreements sit in a long tail (90th percentile for onset: 260 minutes), for example when one device splits or extends the night differently, which is why two wearables have both onset and wake within an hour on 75% of nights rather than nearly all (figure 1b).
7.7 Robustness and an unseen device
Accuracy is similar across reference device types and months (table A4). On the smart rings withheld from development, v2 has onset and wake both within an hour on 67.8% of 41,739 measured nights, slightly above the main result: it generalises to a device whose sleep it was never trained on. Android is five points behind iOS, on 415 Android users against 4,227 on iOS. Fallback windows are much weaker than primary ones (45% against 69% within an hour). The result holds when the reference rule changes and when references identified only indirectly are dropped (table A5).
7.8 Where it fails
On the 13.5% of measured nights more than two hours off (12,907 nights), the causes are concentrated. Assigned by fixed rules in order, the three largest are behaviours the phone cannot disambiguate: phone use in the night read as waking (26%), the phone put down long before sleep beyond what the latency pattern corrects (25%), and a first pickup long after waking (23%). The quiet window itself was wrong on 8%, a fallback window on 4%, and 14% had other causes.
8. What each design choice contributes
This section explains the design with ablations on the main evaluation nights. Each row changes one thing and keeps the rest of the estimator as it is at that stage. These analyses were run after the locked evaluation and are descriptive. Figure 8 summarises the two largest effects: which signals carry the estimate, and the path from v1 to v2.
8.1 The step counter is the backbone
| Signals used | Nights with an estimate | Both within 60 min | More than 2 h off |
|---|---|---|---|
| Screen lock and unlock only | 14.1% | 35.2% | 40.2% |
| App events only | 14.7% | 33.4% | 42.5% |
| Screen and app events, no steps | 26.8% | 45.1% | 32.5% |
| Phone step counter only | 86.5% | 46.4% | 28.9% |
| Step counter plus screen lock | 88.6% | 49.4% | 26.9% |
| Step counter plus screen and app events (as in v2) | 90.4% | 53.3% | 24.6% |
Table 3. Phone signals, with v2’s window rules and a fixed offset (+35 minutes at onset).
Screen and app events alone find a night on only about one occasion in seven (table 3): they exist only when the person uses the app or when the platform records the screen. The step counter is present on nearly every night and carries the estimate; screen and app events sharpen its edges. Every signal added improves both coverage and accuracy.
8.2 Wearable steps quietly mislead development
Adding the wearable’s steps to the phone’s record, as happens by default on a wearable owner’s phone, looks harmless. It is not. With every signal included, coverage falls from 90.4% to 71.2%, because the wearable records small movements through the night that break the quiet window. v1 had been calibrated on data with this mixing, and compensated by ignoring any step bucket under 20 steps. On phone-only data that threshold discards real evidence that the person is up: removing it raises the share of nights within an hour from 44.2% to 53.3%, and coverage from 86.3% to 90.4% (table 4). A phone-only estimator has to be developed on phone-only data; validation sets drawn from wearable owners carry a quiet contamination that tunes the method toward the wrong user.
| Anchor filtering | Nights with an estimate | Both within 60 min | More than 2 h off |
|---|---|---|---|
| Steps below 20 per bucket ignored (v1) | 86.3% | 44.2% | 31.2% |
| All steps kept (as in v2) | 90.4% | 53.3% | 24.6% |
| All steps kept, isolated activity not dropped | 90.3% | 57.1% | 21.2% |
Table 4. Filtering phone activity before finding the quiet window (fixed offsets).
8.3 An inherited filter that costs accuracy
The last row of table 4 removes the isolation filter, which drops activity more than 90 minutes from any other. On phone-only data the filter does more harm than good: without it, the share within an hour rises by 3.8 points. It was inherited from v1 and kept in v2 because it was fixed before this analysis; it is the first candidate for change in the next version.
8.4 Window rules trade coverage for accuracy
| Window rule | Nights with an estimate | Both within 60 min | More than 2 h off |
|---|---|---|---|
| Gap must start 20:00 to 02:00 | 89.2% | 54.4% | 23.2% |
| Gap may start 18:00 to 04:00 (as in v2) | 90.4% | 53.3% | 24.6% |
| No fallback window | 79.8% | 56.6% | 21.4% |
Table 5. Window rules (fixed offsets).
None of these is a free gain (table 5). The fallback adds 10.6 points of coverage on harder nights, which lowers average accuracy. The confidence label (section 7.3) lets a product choose its own operating point instead of the method choosing for it.
8.5 Placement does the most work
| Placing sleep in the window | Both within 60 min | More than 2 h off | Median onset / wake error |
|---|---|---|---|
| Raw window, no offsets | 46.1% | 29.1% | |
| +35 min at onset (v1’s offset) | 53.3% | 24.6% | |
| +35 min at onset, -15 min at wake | 55.7% | 23.6% | |
| Much smaller learned model | 61.3% | 16.5% | 26 / 20 min |
| v2 without personal history | 64.1% | 14.4% | |
| v2 | 65.6% | 13.5% | 24 / 19 min |
Table 6. Placing sleep inside the window (nights with a window, 90.4% of evaluation nights).
Fixed offsets help: starting sleep 35 minutes after the last activity lifts the share within an hour from 46.1% to 53.3%, and ending it 15 minutes before the first activity of the morning lifts it to 55.7%. Learned placement, which uses the latency pattern of section 4.2 and the activity around the window’s edges, lifts it to 65.6% and nearly halves the share more than two hours off, from 23.6% to 13.5%. A much smaller learned model gets most of the way, to 61.3%: most of the gain comes from modelling the latency pattern at all, and the last 4 points from a richer model.
8.6 Personal history helps, a little
Dropping the person’s usual times from v2’s inputs costs 1.5 points (64.1% against 65.6%; table 6). The habit effect in figure 4 is real, but clock time and the night’s own activity carry most of it, because most people’s usual times are close to the population’s. v2 uses history where it exists and works without it from a person’s first night.
8.7 One night in ten has no window
About one night in ten has no usable quiet window, mostly because the phone recorded little that evening or night-time use split the night. Filled nights rest on weaker evidence: onset and wake are both within an hour on 30.6% of them, against 22.9% when such nights are filled from the person’s schedule alone. Including them gives every night an estimate and lowers the all-night share within an hour from 65.6% to 62.2%.
8.8 Duration needs its own model
Used as the duration, the window’s length is 89 minutes too long on average. v2’s duration model brings the mean bias to 4 minutes and the median absolute error from 77 to 39 minutes.
9. Discussion
What the phone supports. A phone, using only what it records anyway, places sleep onset and wake within an hour of a consumer wearable on most nights, and its weekly averages agree within about a quarter of an hour. It does this passively, with no extra sensor, no wearable and nothing for the person to do each night, so it reaches the people a wearable-only programme misses. That supports the features most products build on sleep: regularity, timing, social jetlag, multi-week trends and bedtime guidance. It does not support reporting a single night’s duration as a precise number. Products should present nightly duration as approximate, lean on weekly averages, and use the confidence label to decide which nights to score.
Building on phone-only sleep
- Well supported by these results: sleep timing, bedtime and wake-time guidance, regularity, social jetlag, and weekly averages of timing.
- Approximate: nightly duration (show a range or a weekly average), and week-to-week changes, which track the wearable at r = 0.69.
- Not supported by these results: sleep stages, physiological sleep quality, naps and daytime sleep, and any diagnostic use.
A sleep score for phone-only users should weight timing and regularity, use duration as a weekly average, and treat filled nights as missing.
Lessons beyond this estimator. First, validation data drawn from wearable owners is quietly contaminated by the wearable itself, and settings tuned on it can be wrong for the people a phone-only method exists to serve. Second, the most useful modelling was not a better window but a better account of the minutes around it: when people stop using their phone says a good deal about how long they will take to fall asleep, in a way consistent with the circadian forbidden zone [20, 21].
Comparison with earlier work. Where the numbers are comparable, they are in the same range: median deviations of 43 and 38 minutes for onset and wake against self-report [8], and 89% minute-level accuracy [10], both in small cohorts. Different references and populations make a direct ranking unsound. What this analysis adds is scale, a held-out design fixed in advance, and the coverage, failure and confidence measures that deployment needs.
Next steps. The analysis points to two changes for the next version: drop the isolation filter (section 8.3), and use more of the screen signal where the platform can record it, since every phone signal added improved results (section 8.1).
10. Limitations
- The reference is a consumer wearable, not polysomnography. Wearables agree closely with each other here but share biases relative to laboratory sleep staging.
- The validation set owns wearables; the estimator is for people who do not. Their phones carry similar amounts of signal (a median of 22 phone records a night against 23), but phone-only users have more nights without a usable window (21% filled against 12.5%).
- Single-night duration is weak (ICC 0.45) and regresses toward typical values.
- Filled nights are weaker, with 31% within an hour.
- Most users are on iPhone. Android results rest on 415 people.
- Night sleep only. A night runs from 18:00 to 14:00 and the quiet window must start between 18:00 and 04:00, so daytime sleep after night shifts and naps are not estimated. Nights with a time-zone change were excluded, so travel was not tested.
- Results describe settled estimates, computed after the morning’s data has arrived; an estimate read early the same morning was not evaluated.
- The population is users of apps built on one platform over four summer and autumn months. It is not a population sample, and seasonal effects are untested.
- The design analysis in section 8 was run on the held-out set after the evaluation. It explains choices already made and did not change the estimator, but its numbers are descriptive rather than pre-specified.
11. Responsible use and disclosures
Responsible use. v2 is built to need as little as possible (section 5), but sleep timing is still personal information. It should be shown to the person it describes and used for their benefit, for example in opt-in wellness programmes, in line with Sahha’s Customer Data Use Policy. v2 estimates sleep timing for wellness features; it does not diagnose sleep disorders.
Availability. The design, the signals, the window rules and the inputs of every learned component are described above, with the evidence for each. The data cannot be shared: it is health data from end users of Sahha’s customers.
Competing interests. The analysis was designed, run and written by Sahha staff, and v2 is part of Sahha’s product. To limit the room for choices that favour the result, both candidate estimators were built on development data and the analysis plan was fixed before the held-out evaluation was run once; the choice between the candidates is described in section 6.
Appendix: supporting tables
| Step | Nights | People |
|---|---|---|
| Wearable nights of held-out people in the window | 199,647 | 5,499 |
| Excluded: two or more wearables (used in section 7.6) | 5,405 | 346 |
| Excluded: time-zone change during the night | 3,282 | 811 |
| Excluded: reference sleep shorter than 3 or longer than 13 hours | 9,997 | 2,853 |
| Excluded: nights used in an earlier calibration of v1 | 14,585 | 2,681 |
| Excluded: feature switched off for the account | 356 | 5 |
| Excluded: no phone-origin activity at all | 13,091 | |
| Eligible nights | 152,931 | 4,778 |
| Set aside: unseen-device test (section 7.7) | 47,105 | |
| Main evaluation set | 105,826 | 3,868 |
Table A1. Eligibility for the held-out evaluation. Exclusion counts can overlap.
| Phone went quiet | 2 h+ early | 1 to 2 h early | 0.5 to 1 h early | On time | 0.5 to 1.5 h late | 1.5 h+ late |
|---|---|---|---|---|---|---|
| Before 21:00 | 182 | 117 | 81 | 56 | 49 | |
| 21:00 to 22:00 | 121 | 73 | 48 | 35 | 31 | 30 |
| 22:00 to 23:00 | 94 | 52 | 39 | 28 | 22 | 22 |
| 23:00 to midnight | 39 | 31 | 24 | 18 | 17 | |
| Midnight to 01:00 | 35 | 26 | 20 | 16 | 13 | |
| After 01:00 | 21 | 17 | 11 |
Table A2. Median minutes from last phone activity to sleep onset, by clock time (rows) and relative to the person’s usual time (columns), the values of figure 4. Held-out nights; cells with fewer than 150 nights are blank.
| Minutes 20:00 to 12:00 | Sensitivity | Specificity | Accuracy | κ |
|---|---|---|---|---|
| v2 (measured nights) | 0.923 | 0.868 | 0.892 | 0.781 |
| v1 | 0.953 | 0.763 | 0.844 | 0.692 |
| Population schedule | 0.834 | 0.795 | 0.812 | 0.621 |
| Clock 23:00 to 07:00 | 0.855 | 0.766 | 0.804 | 0.609 |
Table A3. Minute-level agreement, sleep as the positive class, pooled over nights.
| Subgroup (measured nights, all references) | Nights | Both within 60 min | More than 2 h off |
|---|---|---|---|
| Reference: watch | 81,906 | 65.0% | 13.8% |
| Reference: ring | 45,061 | 67.7% | 11.5% |
| Reference: band | 9,903 | 70.6% | 11.1% |
| iOS | 124,702 | 66.7% | 12.5% |
| Android | 12,679 | 61.4% | 16.5% |
| June | 34,565 | 66.1% | 13.0% |
| July | 39,202 | 65.8% | 13.1% |
| August | 39,592 | 66.7% | 12.4% |
| September | 24,022 | 66.5% | 13.0% |
| Primary window | 121,403 | 69.0% | 11.3% |
| Fallback window | 15,978 | 45.5% | 24.6% |
Table A4. Subgroups.
| Sensitivity analysis (measured nights) | Nights | Both within 60 min | More than 2 h off |
|---|---|---|---|
| Main analysis | 95,642 | 65.6% | 13.5% |
| Unseen device only | 41,739 | 67.8% | 11.4% |
| All references | 137,381 | 66.2% | 12.9% |
| References identified only indirectly excluded | 95,049 | 65.6% | 13.5% |
| Reference from every asleep run, not the main episode | 95,642 | 64.2% | 15.5% |
Table A5. Sensitivity of the headline result.
Related research
- The Sahha Research Study: Protocol and Methodology: how Sahha’s earlier mental health studies collected and labelled smartphone data.
- Detecting Stress from Smartphone Activity and Sleep Patterns: a smartphone-only model of stress that uses sleep as one of its inputs.
References
- de Zambotti, M., Cellini, N., Goldstone, A., Colrain, I. M., & Baker, F. C. (2019). Wearable sleep technology in clinical and research settings. Medicine & Science in Sports & Exercise, 51(7), 1538-1557.
- Chinoy, E. D., Cuellar, J. A., Huwa, K. E., et al. (2021). Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep, 44(5), zsaa291.
- Menghini, L., Cellini, N., Goldstone, A., Baker, F. C., & de Zambotti, M. (2021). A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. Sleep, 44(2), zsaa170.
- Viswanath, V. K., Hartogenesis, W., Dilchert, S., et al. (2024). Five million nights: temporal dynamics in human sleep phenotypes. npj Digital Medicine, 7.
- Chen, Z., Lin, M., Chen, F., et al. (2013). Unobtrusive sleep monitoring using smartphones. Proceedings of the 7th International Conference on Pervasive Computing Technologies for Healthcare (PervasiveHealth).
- Abdullah, S., Matthews, M., Murnane, E. L., Gay, G., & Choudhury, T. (2014). Towards circadian computing: “early to bed and early to rise” makes some of us unhealthy and sleep deprived. Proceedings of UbiComp ‘14, 673-684.
- Borger, J. N., Huber, R., & Ghosh, A. (2019). Capturing sleep-wake cycles by using day-to-day smartphone touchscreen interactions. npj Digital Medicine, 2, 73.
- Saeb, S., Cybulski, T. R., Kording, K. P., & Mohr, D. C. (2017). Scalable passive sleep monitoring using mobile phones: opportunities and obstacles. Journal of Medical Internet Research, 19(4), e118.
- Ciman, M., & Wac, K. (2019). Smartphones as sleep duration sensors: validation of the iSenseSleep algorithm. JMIR mHealth and uHealth, 7(5), e11930.
- Cuttone, A., Bækgaard, P., Sekara, V., Jonsson, H., Larsen, J. E., & Lehmann, S. (2017). SensibleSleep: a Bayesian model for learning sleep patterns from smartphone events. PLoS ONE, 12(1), e0169901.
- Knol, L., Ross, M. K., Nagpal, A., et al. (2025). Unobtrusive inference of diurnal rhythms from smartphone data. npj Digital Medicine. doi:10.1038/s41746-025-02254-1.
- Langholm, C., Byun, A. J. S., Mullington, J., & Torous, J. (2023). Monitoring sleep using smartphone data in a population of college students. npj Mental Health Research.
- Kirshenbaum, J. S., Crowley, R. N., Latham, M. D., Pagliaccio, D., Auerbach, R. P., & Allen, N. B. (2025). Comparison of sleep features across smartphone sensors, actigraphy, and diaries among young adults: longitudinal observational study. JMIR Formative Research, 9, e67455.
- Mammen, P. M., Zakaria, C., Molom-Ochir, T., Trivedi, A., Shenoy, P., & Balan, R. (2021). WiSleep: inferring sleep duration at scale using passive WiFi sensing. arXiv:2102.03690.
- Massar, S. A. A., et al. (2021). Trait-like nocturnal sleep behavior identified by combining wearable, phone-use, and self-report data. npj Digital Medicine, 4, 90.
- Windred, D. P., Burns, A. C., Lane, J. M., et al. (2024). Sleep regularity is a stronger predictor of mortality risk than sleep duration. Sleep, 47(1), zsad253.
- Phillips, A. J. K., Clerx, W. M., O’Brien, C. S., et al. (2017). Irregular sleep/wake patterns are associated with poorer academic performance and delayed circadian and sleep/wake timing. Scientific Reports, 7, 3216.
- Wittmann, M., Dinich, J., Merrow, M., & Roenneberg, T. (2006). Social jetlag: misalignment of biological and social time. Chronobiology International, 23(1-2), 497-509.
- Walch, O. J., Cochran, A., & Forger, D. B. (2016). A global quantification of “normal” sleep schedules using smartphone data. Science Advances, 2(5), e1501705.
- Lavie, P. (1986). Ultrashort sleep-waking schedule. III. “Gates” and “forbidden zones” for sleep. Electroencephalography and Clinical Neurophysiology, 63(5), 414-425.
- Strogatz, S. H., Kronauer, R. E., & Czeisler, C. A. (1987). Circadian pacemaker interferes with sleep onset at specific times each day: role in insomnia. American Journal of Physiology, 253(1), R172-R178.
- Exelmans, L., & Van den Bulck, J. (2016). Bedtime mobile phone use and sleep in adults. Social Science & Medicine, 148, 93-101.
- Vogels, E. A. (2020). About one-in-five Americans use a smart watch or fitness tracker. Pew Research Center, 9 January 2020 (survey of 3 to 17 June 2019).
- Pew Research Center (2025). Mobile fact sheet. Smartphone ownership among US adults, 2025.