Sleep Tracker Accuracy: What Your Wearable Measures, and What It Guesses
Medically reviewed for accuracy by Dr. Bruce Rubinowicz, Board-Certified Neurologist & Sleep Medicine Specialist
Quick Answer
Consumer sleep trackers measure movement, heart rate, and sometimes heart rate variability and blood oxygen, then use an algorithm to infer sleep stages; only lab polysomnography (EEG) measures stages directly. Head-to-head testing shows wearables are good at detecting sleep itself, with seven devices identifying sleep at least 93 percent of the time against polysomnography, but poor at catching wakefulness, spotting it only 18 to 54 percent of the time, so quiet tossing gets counted as sleep. A pooled meta-analysis of 24 studies found total sleep time off by about 17 minutes on average, which makes nightly totals and weekly trends useful. Sleep staging is the weak spot: in the same seven-device test, staging accuracy was mixed to often poor, with light sleep systematically overestimated and deep and REM minutes erring in both directions depending on the device. The practical rule is to treat deep sleep and REM numbers as estimates with wide error bars, use trends and consistency rather than absolutes, and run before-and-after experiments like moving your caffeine cutoff. One bad night's sleep score is noise; how you feel across a week beats any single number, and chasing a perfect score, a pattern researchers have named orthosomnia, tends to make nights worse rather than better.
Your watch says you got 42 minutes of deep sleep last night.
Maybe you checked the number this morning, before coffee, and let it set the tone for the day. Plenty of people do, and the ritual has a flaw baked into it: your watch never measured your deep sleep. It can't. Sleep stages are defined by brain activity, and there's no electrode on your wrist. What the watch recorded overnight was movement and heart rate — maybe heart rate variability and blood oxygen too — and then an algorithm guessed which stages would best explain those signals. Sometimes the guess is decent. Often it isn't.
This isn't an anti-tracker piece. The research record says consumer wearables have gotten good at one job and stayed shaky at another, and once you know which numbers on the morning screen are measurements and which are inferences, all of them get more useful. You know how hard to lean on each.
What a wrist can see
A consumer wearable carries a short list of sensors. An accelerometer tracks movement. An optical sensor shines light into your skin and reads your pulse, which gives heart rate and heart rate variability; newer models add skin temperature and blood oxygen. That's the whole list, and nothing on it touches the brain.
Sleep stages are a brain thing. Light sleep, deep sleep, and REM are categories scored from electrical brain activity, eye movements, and muscle tone, and the instrument that captures those signals is polysomnography — the overnight lab study where electrodes go on your scalp and face (the whole paste-in-your-hair production) and a trained scorer labels every 30-second window of the night. Polysomnography, PSG for short, is the only direct measure of sleep stages there is. Everything a wrist device says about stages is a statistical translation: body still, heart rate down, variability in a certain pattern, so probably this stage.
That translation can be tested, and sleep science built a standard playbook for testing it. A 2021 methods paper in SLEEP (Menghini and colleagues) laid out the now widely used framework: put the wearable and PSG on the same sleeper on the same night, then compare them window by window across the whole night, with agreement plots that show whether a device errs consistently or wildly. Before that framework, accuracy claims weren't comparable across studies. After it, they were.
The report card: good at sleep, bad at wake
The clearest single test is a 2021 study in SLEEP (Chinoy and colleagues) that ran seven consumer sleep trackers simultaneously against PSG in 34 healthy adults over three lab nights, one of them deliberately disrupted. On telling sleep from wake, the devices did well: every one identified sleep at least 93 percent of the time, and most matched or beat actigraphy, the research-grade wrist monitor sleep science leaned on for decades. On that axis, the thing on your wrist has caught up with the professional tool.
The weakness sits on the other side of the ledger. When participants were awake in bed, the devices caught it only 18 to 54 percent of the time, depending on the model — the best device in the study missed roughly half the wake time, and the worst missed more than four fifths of it. The reason is mechanical. Lie still in the dark with a calm pulse, running through tomorrow's list, and you look asleep to an accelerometer and a heart rate sensor. There's no motion signature that cleanly separates quiet rest from light sleep, so the tracker rounds you up.
An expert panel convened by the Sleep Research Society reached the same verdict when it reviewed the field in a 2024 state-of-the-science paper in SLEEP (de Zambotti and colleagues): consumer wearables now beat old-style actigraphy at assessing sleep against PSG, and their clearest named limitation is misclassifying wakefulness during the night. The panel also noted that accuracy isn't uniform across people — factors like skin tone can degrade the optical signal — which is one more reason to hold any individual reading loosely.
How big is the error in practice? A meta-analysis in the Journal of Clinical Sleep Medicine (Lee and colleagues, 2025) pooled 24 studies covering 798 sleepers across a spread of brands. On average, devices underestimated total sleep time by about 17 minutes and overestimated overnight wake time by about 13 minutes, with wide variation between devices and studies. The direction isn't even consistent across the literature — the pooled average leaned toward undercounting sleep, while the missed-wake problem in the seven-device lab data pushes the other way on restless nights. A sensible reading of all of it: treat any single night's sleep total as accurate to within roughly half an hour, either direction.
| What the app shows | Where it comes from | Agreement with lab PSG | How to treat it |
|---|---|---|---|
| Total sleep time | Movement plus heart rate | Good; pooled average error about 17 minutes (Lee et al., 2025) | Trust the weekly trend; allow half an hour of slack on any one night |
| Sleep vs. wake | Same sensors | Sleep detected 93%+ of the time; wake caught only 18-54% of the time (Chinoy et al., 2021) | Assume quiet wakefulness was undercounted |
| Sleep timing and consistency | Clock time of detected sleep | The strongest suit of the whole category | The most trustworthy screen in the app |
| Deep sleep minutes | Algorithmic inference from heart rate, variability, and movement | Mixed to poor; three of six staging devices overestimated it (Chinoy et al., 2021) | An estimate with wide error bars; compare only to your own baseline |
| REM minutes | Same inference | Mixed; three of six devices underestimated it | Same treatment |
| Sleep score | Proprietary formula | Never validated as a composite | Ignore daily swings; glance at the monthly drift |
The deep sleep number is an estimate in a lab coat
Staging is where the honest answer gets uncomfortable for the morning ritual. In the seven-device study, the authors described stage detection as mixed and often poor. All six devices that scored stages overestimated light sleep. Three overestimated deep sleep, three underestimated REM, and misclassified windows tended to get dumped into the light sleep bucket. The night-to-night scatter was large, too — a device could land near the truth one night and miss badly the next, on the same wrist.
There's a second, quieter problem: the algorithm doing the guessing is proprietary. Manufacturers don't publish how stages are inferred, and a firmware update can change the scoring without notice, so your deep sleep numbers from last winter and this summer may not even come from the same math. Validation studies test a snapshot of a device. Your wrist runs whatever shipped last month.
So a 20-minute swing in deep sleep between last night and tonight is usually just noise in the inference. Averaged over weeks, the number can still drift the right way when your sleep really changes, and that's about all it's for.
When the score starts running your nights
There's a failure mode beyond misreading the data, and sleep clinicians have watched it walk through the door. A case series in the Journal of Clinical Sleep Medicine (Baron and colleagues, 2017) described people who arrived at sleep clinics organized entirely around their tracker output, and coined a name for the pattern: orthosomnia, the perfectionist pursuit of ideal sleep numbers. One man chased eight hours of what his tracker labeled deep sleep and stretched his time in bed to farm the minutes. Another trusted her wristband over an overnight lab study that contradicted it — why did the device say she slept badly when the electrodes said otherwise? A third fixated on early-morning wakeups only his tracker could see.
These were people troubleshooting, to the minute, a number the device can't measure to the minute. And the main coping move, spending more time in bed to raise the score, runs exactly backward: more time in bed usually means more of it spent lying awake, which makes nights lighter and choppier, which lowers the score, which invites more time in bed. The tracker becomes the stressor it was supposed to monitor.
Part of what feeds the loop is that trackers surface things that were always there. Brief nighttime wakeups are a normal feature of sleep architecture, and most of them used to vanish from memory by morning — we walked through the mechanics in why you wake up at 3 AM. A device that itemizes every one of them can convert normal into alarming. If seeing the wakeups listed makes you fret about them, that's a reason to hide the screen, not a reason to fix sleep that may not be broken.
What the thing is good for
Used well, a tracker earns its spot on your wrist three ways, and all three play to what the sensors measure.
First, trends. A device that's off by 20 minutes tends to be off by a similar amount in a similar direction night after night, so the error largely subtracts out when you compare you to you. Six weeks of your own totals drifting down is a real signal even if no single night's number is exact.
Second, consistency. Timing is what a wearable measures best, because detecting when a sleep bout started and ended is far easier than grading its interior. A steady sleep window and a steady wake time are two of the highest-leverage sleep behaviors there are, and the tracker audits both for free. If you want the wake side of that equation to pull its weight, pair the data with the first-90-minutes morning framework, which is built around a consistent get-up time.
Third, before-and-after experiments — the most underused feature of the whole category. Pick one variable, change it for two weeks, and compare your own trend lines. Move your caffeine cutoff from mid-afternoon to mid-morning and watch total sleep time and overnight wake estimates. Shift alcohol earlier in the evening, or drop it for a stretch, and do the same. The absolute numbers can be off while the direction of change is still informative, which is all an experiment needs from them.
When to ignore the score
One bad night's score is noise twice over: the algorithm's single-night error is large, and normal sleep varies night to night in ways that don't need explaining. A rough score after a stressful evening tells you nothing you didn't already know. A rough score after a night that felt fine is more likely a staging artifact than a hidden problem.
So run the hierarchy in order. How you feel across a week outranks any number. The weekly trend outranks any single night. Any single night's total outranks its stage breakdown, because sleep versus wake is measured while stages are inferred. When the score and your body disagree for days on end, believe your body — and if you feel persistently unrefreshed despite enough time in bed, that's a conversation for a clinician and possibly a lab study, which measures what a wrist can't.
Think of the tracker the way you think of a bathroom scale. Nobody weighs themselves to learn the ounce; the scale is for the trend, and if you're stepping on it five times a day, the scale has become the problem. The watch knows roughly when you slept and for how long, and it's good at that.
It guesses the rest.
Sources
Frequently Asked Questions
How accurate are sleep trackers compared to a sleep study?
They're good at detecting sleep and weak at detecting wake. In a seven-device test against lab polysomnography, every tracker correctly identified sleep at least 93 percent of the time, and a pooled meta-analysis put the average total sleep time error around 17 minutes. The same devices caught wakefulness only 18 to 54 percent of the time, though, and sleep staging accuracy was mixed to poor. A lab study is still the only direct measure of sleep stages.
Is the deep sleep number on my watch accurate?
It's an estimate, and often a rough one. Deep sleep is defined by brain waves; your watch reads movement and heart rate, then guesses. In head-to-head testing, devices systematically overestimated light sleep while deep and REM errors ran in both directions depending on the model. Compare the number to your own baseline over weeks, never to a friend's watch or an ideal target.
Why does my tracker say I was asleep when I was lying awake?
Because lying still with a calm heartbeat looks identical to sleep from the wrist. Validation studies found consumer devices detect wake only 18 to 54 percent of the time, so quiet wakefulness gets rounded up to sleep. That's why a restless night often scores better than it felt — the sensors missed the part where you were awake.
What is a good sleep score?
Nobody can tell you. The score is a proprietary formula that differs by brand and can change with a software update, and no study has validated the composite score itself against lab measurement. Ignore single-night swings and watch the multi-week drift instead. A score that trends steady while you feel fine is all the reassurance the number can give.
Should I stop wearing my sleep tracker?
Keep it if it's helping, take it off if it's stressing you. A tracker that nudges you toward a consistent schedule and lets you run before-and-after experiments is earning its keep. But if checking the score adds stress, or you catch yourself staying in bed longer to farm minutes, shelve it for two weeks — researchers have documented that chasing perfect tracker numbers can make nights worse. How you feel across a week is the better instrument.

From Risachi
Sleep Collection Gummies
20mg CBN + 15mg CBG + 500mg Montmorency tart cherry per gummy. Melatonin-free.
- Melatonin-free — no morning fog
- Third-party lab tested, COA published
- Vegan, non-habit forming
$39.9924 Gummies per bottle
60-day money-back guarantee · Bundle savings available
Reading List
Get the next breakdown before anyone else
New sleep-science articles and lab-testing explainers, reviewed by a board-certified sleep-medicine physician. No daily email.
Educational content only — not medical advice. Unsubscribe anytime.