What Oura, Whoop, and Eight Sleep Actually Measure, and Where the Numbers Stop Being Reliable

You wake up feeling reasonably good. You reach for your phone. Your recovery score is 34 percent.

Or the reverse, which is worse. You wake up wrecked, the kind of tired that sits behind your eyes, and the app says you are at 91 percent and cleared for a hard session. Either way a small argument starts, and by the time you are on the train you have decided that either your body is wrong or your ring is.

Almost everyone wearing one of these devices in New York has had that morning. What almost nobody has is a framework for resolving it, because the conversation has collapsed into two useless positions. One says the data is objective truth and your feelings are noise. The other says it is a gimmick. Neither survives contact with the validation literature.

‍The distinction that resolves most of this is simple, and once you have it you will not be able to unsee it. The sensors and the scores are two completely different things, with two completely different evidence bases.

The sensors are genuinely good

Start with what these devices physically detect, because it is less mysterious than it seems.

‍A ring or a wrist strap uses photoplethysmography, which means shining light into your skin and measuring how much comes back. Blood absorbs light, and the volume of blood under the sensor changes with every heartbeat, so the returning signal traces out your pulse. From the spacing between those pulses you get heart rate and, with more processing, heart rate variability. Add an accelerometer for movement and a thermistor for skin temperature and you have the full input set for most wrist and finger wearables.

‍Eight Sleep belongs in its own category. It sits under you rather than on you, reading pressure changes through the mattress, which means it infers cardiac and respiratory activity from the mechanical vibration your body transmits into the bed. It is also the only one of the three that intervenes rather than only measures, since thermal regulation actively changes the sleep environment.

On the core measurements, the independent evidence is reassuring. A 2025 validation study in Physiological Reports put five devices against a gold-standard ECG reference across 536 nights, and accuracy for overnight resting heart rate was very high, with the Oura generations showing roughly two percent mean error. Heart rate variability was noisier but still respectable, around six to eight percent error for the better performers. For sleep versus wake detection, a study of the Oura Gen 3 against full polysomnography across 96 participants and more than 420,000 scored epochs found overall accuracy near 92 percent.

Those are independent laboratory comparisons, not marketing numbers. The hardware does what it claims.

Where it gets more complicated

‍Now the part that does not make it into the product page.

‍In that same polysomnography study, the device correctly identified sleep about 94 percent of the time but correctly identified wakefulness only about 73 percent of the time. Its predictive value for wake was around 67 percent. In plain terms, these devices are excellent at knowing when you are asleep and comparatively poor at knowing when you are lying still and awake.

‍That asymmetry explains a lot of frustrating mornings. If you spent forty minutes at three in the morning awake but not moving, thinking about a deadline, the device likely scored much of that as sleep. You know you were awake. The ring reports a decent night. Your body is not lying to you.

‍Sleep staging is looser still. Distinguishing light sleep, deep sleep, and REM from a peripheral pulse signal is hard, and stage-level accuracy in that study ranged from about 76 percent for light sleep to about 91 percent for REM. A systematic review of Whoop found the same shape: acceptable for two-stage sleep and heart rate, with room for improvement on four-stage sleep and HRV.

Here is the context that makes this fair rather than damning. When two trained technicians independently score the same night of laboratory sleep data, they agree roughly 83 percent of the time. The gold standard is not a perfect ruler either. A consumer device landing in the high seventies on a task where trained humans hit the low eighties is doing something real.

The scores are a different category entirely

‍Everything above concerns measurement. Almost nothing you actually look at in the morning is a measurement.

‍Readiness, Recovery, and Sleep Score are composites. They take several measured inputs, weight them through a proprietary algorithm, compare them against your personal baseline, and output a single number. The authors of that 2025 validation study made the point directly, noting that there is little transparency into what actually drives the readiness and recovery scores the end user sees, and that these algorithms get updated periodically. The number can change meaning underneath you without any change in your physiology.

So the situation is this. The heart rate is validated. The HRV is reasonably validated. The sleep-versus-wake call is validated. The number that determines whether you feel good about your morning is not independently validated at all.

‍That is not a scandal, just a normal consequence of proprietary product development. But it does mean treating a recovery score as an objective verdict is a category error, and the people most likely to make it are the ones taking the whole project most seriously.

Old school and new school are measuring the same thing badly in different directions

The instinct is to frame this as data versus intuition, with one of them winning. That framing is wrong, and dropping it is the most useful thing in this article.

‍Your subjective sense of how you feel is also data. It is a genuine physiological readout integrating interoceptive signals, autonomic state, sleep pressure, and mood. It has a well-known error profile, being heavily influenced by expectation, by what happened yesterday, by caffeine, and by what you have already decided about the day.

‍The device has a different error profile. It is poor at single-night verdicts and quite good at multi-week trends, because random error averages out over time while systematic drift shows up clearly. It is also bad at attribution, meaning it can tell you something shifted but not why. Your felt sense is often better at the why and worse at the magnitude.

‍Those two error profiles barely overlap, which is exactly why using both beats using either. The device is a trend line. Your body is a signal. Neither is a verdict.

‍The practical version: stop reading the daily number and start reading the fourteen-day shape. If your overnight heart rate has been climbing for ten days and your HRV has been drifting down, that is worth taking seriously regardless of how you feel, because that pattern is well within what the sensors measure reliably. If your score is 41 on a Tuesday and you feel fine, that is within the noise, and reorganizing your day around it is reading precision into a number that does not have it.

‍There is a documented version of getting this wrong. Sleep clinicians have described a pattern where people become anxious about optimizing their sleep data, and the anxiety itself degrades their sleep. In a city that already runs on measurement and performance, this is not a hypothetical risk. It is the most common way a useful tool becomes a harmful one.

What the numbers are actually pointing at

Here is where this connects to the thing we do.

‍A recovery score is not a property of the device. It is a compressed readout of your autonomic state, your inflammatory load, your sleep architecture, and your body's current capacity to repair itself. Those are real physiological variables, and they are the things worth working on. The score is a window onto them, sometimes a smudged one.

You cannot train a number. You can change the physiology the number is reading.

Hyperbaric Oxygen Therapy is a systemic modality that influences the human body on cellular and physiological level. The variables it touches are the same ones sitting underneath your readiness score. Inflammatory signaling, mitochondrial function, vascular delivery, and autonomic regulation are the substrate of what the wearable is estimating, and the relationship between hyperbaric therapy and mitochondrial behavior is one part of that picture.

‍This is also where wearables become legitimately useful rather than decorative. If you are going to try a systemic intervention, having several weeks of baseline HRV and overnight heart rate data before you start is genuinely valuable, because those are the measurements the devices do well. Watching the trend across a course of sessions is more informative than asking yourself whether you feel different, which is exactly the question your subjective sense is worst at answering. We see this with people recovering from post-viral states, where improvement is often gradual enough to be invisible day to day and obvious in a trend line, and with athletes managing training load, where the whole question is whether recovery is keeping pace with demand.

Outcomes vary and the body is not one switch. No honest reading of the data would tell you a wearable predicts who responds to anything. What it can do is give you a more reliable answer than memory about whether something changed.

The money question

‍There is a financial contrast worth making, and it does not favor us in the way you might expect.

A serious wearable setup runs a few hundred dollars plus a subscription, and an Eight Sleep configuration runs into the thousands. People buy these readily, because the cost is spread out and the object is pleasant to own. The same people will then hesitate over a course of sessions that might actually move the variables the device is measuring.

‍That is worth noticing. The wearable tells you the trend is bad. It does not change the trend, and it has no mechanism to. If you have tracked a declining recovery trend for six months and bought better tracking each time, you have paid for increasingly precise measurement of a problem you have not addressed. At some point the useful spend shifts from measuring to intervening, and knowing where that point sits may be the most valuable thing the data gives you.

The point of the number

‍The devices are better than the skeptics think and less authoritative than the enthusiasts think. Sensors: trustworthy. Trends: trustworthy. Daily verdicts: not really. Proprietary scores: unverifiable by anyone outside the company.

‍Use them as instruments rather than judges. Read the multi-week shape and ignore the daily fluctuation. Let a bad number prompt a question rather than dictate a decision. And if the trend has been flat for a long stretch while you have been doing everything correctly, that usually means the input needs to change rather than the tracking.

‍If you want to talk through where a systemic modality might fit alongside what you are already measuring, our overview of HBOT is the place to start, and you can book a session or a conversation here. Before you commit to anything, our breakdown of what HBOT costs in New York and why lays out the economics honestly.

The goal was never a better score. It was waking up and recognizing the person you used to be before you needed a device to tell you how you were doing.

Frequently Asked Questions

Previous
Previous

Hard Chamber vs Soft Chamber: What the Difference Actually Means for Your Body

Next
Next

What Hyperbaric Oxygen Therapy Costs in New York City, and What Drives the Price